宜家爬虫报错UnboundLocalError:局部变量title未赋值即引用
问题分析与修复
错误根源
你碰到的UnboundLocalError核心原因是:title变量仅在for循环内部赋值。如果页面加载后找不到任何per-product-link-wrapper元素(也就是cards是空列表),循环根本不会执行,title就从未被定义,到后面写入JSON文件的代码行自然触发报错。
另外代码里还有个隐藏bug:最后一行products_info.append(products_info)会把列表自身嵌套进去,生成递归错误的数据结构,这绝对不是你想要的结果。
修复后的代码
调整逻辑,把title的赋值提到循环外(页面标题是整个分类的,没必要在循环里重复获取),同时处理空数据场景,修正列表自嵌套的错误:
def get_info_from_url(browser, sub_category, url_main): print(sub_category) products_info = [] browser.get(url_main + "?page=200") time.sleep(10) # 提前获取并处理标题,彻底避免循环不执行时变量未定义 title = browser.title.replace(" - IKEA", '') cards = browser.find_elements(By.CLASS_NAME,'per-product-link-wrapper') for card in cards : product_name = card.find_element(By.CLASS_NAME, "withprice-title").text product_desc = card.find_element(By.CLASS_NAME, 'withprice-commit').text product_price = card.find_element(By.CLASS_NAME,'withprice-price').text print(product_name, product_desc ,product_price) product_url = card.get_attribute('href') product_code = product_url.split("-")[-1].strip("s/") prod = { 'sub_category': title, 'product_info': [{ 'name': product_name, 'desc': product_desc, 'price': product_price, 'product url': product_url, 'product code': product_code }] } products_info.append(prod) # 处理空数据,避免写入空文件或报错 if not products_info: print(f"警告:{sub_category}分类下未找到任何产品") return products_info # 确保保存目录存在,避免写入时因目录缺失报错 import os save_dir = f'JSON/{sub_category}' if not os.path.exists(save_dir): os.makedirs(save_dir) with open(f'{save_dir}/{title}.json', 'w', encoding='utf-8') as outfile: json.dump(products_info, outfile, ensure_ascii=False, indent=2) # 移除错误的列表自嵌套代码 return products_info
额外优化建议
- 替换固定
time.sleep(10)为显式等待,提升爬虫稳定性和效率:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 替换原time.sleep(10) WebDriverWait(browser, 20).until( EC.presence_of_element_located((By.CLASS_NAME, 'per-product-link-wrapper')) )
- 给单个
card.find_element加try-except捕获异常,避免单个产品爬取失败导致整个分类的爬取流程中断。
内容的提问来源于stack exchange,提问作者IzTuugii
相关产品推荐
相关产品推荐

