You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫初学者用BeautifulSoup提取商品标题返回None问题求助

问题原因

  • 目标站点基于Angular开发,商品标题属于JS动态渲染的内容:requests库只能拿到未执行JS的原始静态HTML,此时你要查找的h4.ng-binding元素还未生成,所以查找结果为None。
  • 你可以打印site.content查看实际拿到的源码,确认不存在对应元素。

解决方案

方案1:使用支持JS渲染的爬取工具

使用selenium、playwright等模拟浏览器的工具,等待页面JS执行完成后再提取元素,selenium示例代码如下:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options

URL = "https://pauta.com.br/produto/31394"
chrome_options = Options()
# 无头模式后台运行,不需要弹出浏览器窗口
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
driver.get(URL)
# 隐式等待最长10秒,等元素加载完成
driver.implicitly_wait(10)

title = driver.find_element(By.CSS_SELECTOR, 'h4.ng-binding').text
print(title)

driver.quit()

方案2:直接请求数据接口

打开浏览器开发者工具的「网络」面板,筛选Fetch/XHR请求,找到加载商品数据的后端接口,直接请求接口获取JSON格式的商品数据,该方案效率远高于模拟浏览器,也无需处理渲染逻辑。

内容的提问来源于stack exchange,提问作者puppetmaster12039

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 16:45:04