You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup爬取1mg页面遇AttributeError错误求助

问题分析与解决建议

核心问题1:代码顺序错误触发异常

你当前代码先执行print(product_title_element.text),再判断元素是否为None。只要product_title_element是None,这行代码直接抛出AttributeError,后续判断逻辑根本无法执行。必须先判断元素是否存在,再调用text属性。

核心问题2:网站反爬拦截导致返回非目标页面

多数电商网站会检测请求来源,直接用requests.get()发送请求会被识别为爬虫,返回的HTML内容和浏览器中看到的不一致,自然找不到目标元素。需要添加请求头模拟浏览器访问。

修复后的代码示例

import requests
from bs4 import BeautifulSoup

url = "https://www.1mg.com/otc/durex-invisible-super-ultra-thin-condom-otc593647"

# 添加请求头模拟浏览器
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

response = requests.get(url, headers=headers)

# 先检查请求是否成功
if response.status_code != 200:
    print(f"请求失败,状态码:{response.status_code}")
else:
    soup = BeautifulSoup(response.content, "html.parser")
    product_title_element = soup.find("h1", {'class': 'ProductTitle__product-title___3QMYH'})
    
    # 先判断元素存在性再处理文本
    if product_title_element is not None:
        product_title = product_title_element.text.strip()
        print(product_title)
    else:
        print("未找到目标元素,请检查class名或页面内容")

额外排查建议

  • 打印response.text查看返回内容是否和浏览器页面一致,如果是空白或含反爬提示,需进一步处理反爬(比如添加更多请求头、使用代理等)。
  • 再次核对目标元素的class名,部分网站的class是动态生成的,刷新页面后可能变化。

内容的提问来源于stack exchange,提问作者jhone aish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 20:18:19