如何爬取早餐菜单?BeautifulSoup遇NoneType属性错误求助
问题:提取早餐菜单的Fruit Variety内容
想要提取目标早餐菜单页面中的Fruit Variety内容,对应菜单表格截图:
尝试的代码
参考教程编写了如下代码:
import requests from bs4 import BeautifulSoup url ="https://dcsd.nutrislice.com/menu/meadow-view/breakfast/2023-04-14" doc =requests.get(url).content tags =BeautifulSoup(doc,'html.parser') # print(tags.prettify()) parent = tags.find("body").find("ul") text = list(parent.descendants) print(text)
遇到的错误
运行代码后抛出错误:
Traceback (most recent call last): File "C:\Users\User\PycharmProjects\Data_Science\get_content.py", line 8, in <module> text = list(parent.descendants) AttributeError: 'NoneType' object has no attribute 'descendants'
推测原因是页面数据由JavaScript渲染,导致直接用requests获取的静态HTML中没有目标内容。
解决提示
- 使用支持JavaScript渲染的工具,比如
selenium:通过浏览器驱动加载页面,等待JS渲染完成后再抓取DOM内容。示例代码大致结构:from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup url = "https://dcsd.nutrislice.com/menu/meadow-view/breakfast/2023-04-14" driver = webdriver.Chrome() # 需要提前安装对应浏览器驱动并配置环境 driver.get(url) # 等待目标元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//*[contains(text(), 'Fruit Variety')]")) ) # 获取渲染后的页面源码 doc = driver.page_source soup = BeautifulSoup(doc, 'html.parser') # 精准定位Fruit Variety对应的内容 fruit_variety_item = soup.find(text='Fruit Variety').find_next_sibling() print(fruit_variety_item.get_text(strip=True)) driver.quit() - 检查页面是否有公开API接口:打开浏览器开发者工具的Network标签,刷新页面后筛选XHR/Fetch请求,大概率能找到直接返回菜单数据的API,直接请求API获取JSON格式数据,比渲染页面更高效。
- 放弃模糊的元素定位方式:即使页面渲染正常,也不要用
find("body").find("ul")这种宽泛的选择器,应该根据元素的类名、ID或文本关联的XPath来精准定位目标元素,降低找不到元素的概率。
内容的提问来源于stack exchange,提问作者neural science
相关产品推荐
相关产品推荐

