You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup爬虫continue失效,如何跳过无效页面继续循环?

问题原因

你的判断逻辑存在两个核心错误,导致无法跳过无效页面:

  • 判定条件写错了对象:minitabSection是目标页面内div标签的class属性值,永远不会出现在你拼接生成的URL字符串里,这个if判断永远不会成立,continue逻辑根本不会触发。
  • 缺少异常兜底:无效页面、请求失败、页面结构变动都会导致后续元素提取时抛出NoneType错误、索引越界错误,直接终止整个循环。
修复步骤
  1. 移除无效的URL字符串判断,改为直接判断BeautifulSoup查找到的div_techspec元素是否存在——无效页面找不到这个div,返回值为None,直接跳过当前循环即可。
  2. 给网络请求加异常捕获,兜住超时、连接错误、404/500类HTTP错误,避免请求阶段报错卡断循环。
  3. 移除重复的find_all调用(原有代码四次调用完全是查找同一组li标签,属于冗余代码),同时增加li标签数量校验,避免索引越界报错。
  4. 删掉没用的from unittest import skip导入,把JS风格的//注释改成Python支持的#注释。
修复后可运行代码
import requests
from bs4 import BeautifulSoup
from csv import writer
import openpyxl

wb = openpyxl.load_workbook('Book3.xlsx')
ws = wb.active

with open('mtbs.csv', 'w', encoding='utf8', newline='') as f_output:
    csv_output = writer(f_output)
    header = ['Code','Product Description']
    csv_output.writerow(header)
    
    for row in ws.iter_rows(min_row=1, min_col=1, max_col=1, values_only=True):
        product_code = row[0]
        # 跳过Excel里的空行
        if not product_code:
            continue
        url = f"https://www.radwell.com/en-US/Buy/MITSUBISHI/MITSUBISHI/{product_code}"
        print(f"正在采集: {url}")
        try:
            # 加10秒超时,避免请求无限卡死
            req_page = requests.get(url, timeout=10)
            # 遇到4xx/5xx类错误状态码直接抛错,进入异常跳过分支
            req_page.raise_for_status()
        except Exception as e:
            print(f"请求失败,跳过编码{product_code},错误信息:{str(e)}")
            continue

        soup = BeautifulSoup(req_page.content, 'html.parser')
        div_techspec = soup.find('div', class_="minitabSection")
        # 找不到目标div说明是无效产品页,直接跳过
        if not div_techspec:
            print(f"编码{product_code}对应页面不存在,跳过")
            continue

        li_list = div_techspec.find_all('li')
        # 校验页面li标签数量足够,避免索引越界报错
        if len(li_list) < 4:
            print(f"编码{product_code}页面结构异常,字段不足,跳过")
            continue

        info = [li_list[0].text, li_list[1].text, li_list[2].text, li_list[3].text]
        csv_output.writerow(info)

内容的提问来源于stack exchange,提问作者user19380275

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 00:36:17