Python网页爬取:Zomato餐馆名称获取返回NoneType问题求助
爬虫返回NoneType的问题排查与Zomato餐馆名称爬取方案
问题根源分析
- URL错误:你代码里写的是亚马逊的URL,和实际要爬的Zomato完全不相关,这是最直接的问题。
- 元素定位失效:即使换成Zomato的URL,也可能因为以下原因拿不到元素:
- Zomato部分内容是JavaScript动态加载的,
requests只能获取静态HTML,看不到动态渲染后的元素。 - 页面类名可能随版本更新变化,你浏览器检查元素看到的类名,在静态HTML里可能不存在。
- 未处理反爬机制,比如缺少必要请求头(如语言头、Cookie),导致服务器返回的内容不完整。
- Zomato部分内容是JavaScript动态加载的,
解决方案
1. 修正URL并尝试静态爬取
先把URL替换为Zomato的目标页面(比如某城市的餐馆列表页),再检查静态HTML里是否包含目标内容:
from bs4 import BeautifulSoup import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36", "Accept-Language": "en-US,en;q=0.9" # 避免本地化内容差异 } # 替换为Zomato目标页面URL,示例为新德里餐馆列表 url = "https://www.zomato.com/new-delhi/restaurants" response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 用模糊匹配类名的方式查找餐馆名称 restaurant_names = soup.find_all("h4", class_=lambda x: x and "sc-1hp8d8a-0" in x) if not restaurant_names: print("静态HTML中未找到餐馆名称,需处理动态加载内容") else: for name in restaurant_names: print(name.get_text(strip=True))
2. 处理动态加载内容(推荐)
如果静态HTML里没有目标内容,说明是JS动态渲染的,用selenium模拟浏览器加载页面:
先安装依赖:
pip install selenium
下载对应浏览器的驱动(如ChromeDriver,需和浏览器版本匹配),然后运行代码:
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time options = Options() options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") options.add_argument("--headless=new") # 无头模式,不显示浏览器窗口 driver = webdriver.Chrome(options=options) url = "https://www.zomato.com/new-delhi/restaurants" driver.get(url) time.sleep(3) # 等待页面加载完成 soup = BeautifulSoup(driver.page_source, "html.parser") # 需自行验证当前Zomato页面的餐馆名称类名,示例为常见的组合类名 restaurant_names = soup.find_all("h4", class_="sc-1hp8d8a-0 sc-1hp8d8a-2") for name in restaurant_names: print(name.get_text(strip=True)) driver.quit()
3. 关键注意事项
- 类名验证:每次爬取前,用浏览器「查看网页源代码」(而非检查元素)确认类名是否存在于静态HTML中,避免被动态DOM误导。
- 反爬处理:Zomato有反爬机制,频繁请求可能被封IP,建议添加请求间隔、使用代理IP,必要时模拟登录。
- 选择器优化:不要过度依赖自动生成的长类名(如
sc-1hp8d8a-0.sc-lffWgi.flnmvC),这类类名易变化,优先用父元素+标签名的稳定结构。
内容的提问来源于stack exchange,提问作者Tanishqa Garg
相关产品推荐
相关产品推荐

