You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup仅抓取Yelp餐厅名称,排除分页页码?

如何用BeautifulSoup从Yelp抓取餐厅名称并排除分页页码?

问题描述

我正在尝试从Yelp的旧金山餐厅搜索页面抓取餐厅名称,但输出结果中同时包含了分页页码。请问如何使用BeautifulSoup仅定位目标餐厅名称?附上检查元素截图想了解如何定位名称属性,当前代码如下:

import requests
from bs4 import BeautifulSoup as bs
url = "https://www.yelp.com/search?cflt=restaurants&find_loc=San+Francisco%2C+CA"
yelp_r = requests.get(url)
yelp_soup = bs(yelp_r.text, "html.parser")
# print(yelp_soup.prettify())
for name in yelp_soup.find_all("a", {"class": "lemon--a__373c0__IEZFH link__373c0__1G70M link-color--inherit__373c0__3dzpk link-size--inherit__373c0__1VFlE"}):
    print(name.text)

解决方案

问题出在你使用的class是Yelp页面中多个元素共用的——分页的页码链接也使用了这个类,所以会被一起抓取。要精准定位餐厅名称,我们可以通过缩小元素范围来实现,这里提供两种可靠的方法:

方法1:通过餐厅卡片父容器过滤

观察页面HTML结构,餐厅名称会嵌套在专属的卡片容器内,而分页元素在完全独立的容器里。我们先定位所有餐厅卡片,再从中提取名称:

import requests
from bs4 import BeautifulSoup as bs

url = "https://www.yelp.com/search?cflt=restaurants&find_loc=San+Francisco%2C+CA"
yelp_r = requests.get(url)
yelp_soup = bs(yelp_r.text, "html.parser")

# 定位所有餐厅卡片容器(class可能随Yelp更新,需根据当前页面调整)
restaurant_cards = yelp_soup.find_all("div", class_="lemon--div__373c0__1mboc container__373c0__2T9jB hoverable__373c0__1VFlE arrange__373c0__2C9bH border-color--default__373c0__3-ifU")

for card in restaurant_cards:
    # 在卡片内查找名称链接
    name_link = card.find("a", class_="lemon--a__373c0__IEZFH link__373c0__1G70M link-color--inherit__373c0__3dzpk link-size--inherit__373c0__1VFlE")
    if name_link:
        print(name_link.text.strip())

方法2:使用精准的CSS选择器

如果检查元素时发现餐厅名称的<a>标签有独特的层级关系(比如嵌套在<h3>标签下),可以直接用CSS选择器定位:

import requests
from bs4 import BeautifulSoup as bs

url = "https://www.yelp.com/search?cflt=restaurants&find_loc=San+Francisco%2C+CA"
yelp_r = requests.get(url)
yelp_soup = bs(yelp_r.text, "html.parser")

# 选择h3标签下的目标链接,避开分页元素
for name in yelp_soup.select("h3 > a.lemon--a__373c0__IEZFH.link__373c0__1G70M"):
    print(name.text.strip())

实用小技巧

Yelp的页面结构偶尔会更新,若之后class失效,你可以重新用浏览器检查元素:

  • 右键点击餐厅名称 → 选择「检查」
  • 观察该<a>标签的父元素是否有独特的class/属性
  • 基于父元素缩小抓取范围,就能轻松避开分页链接

内容的提问来源于stack exchange,提问作者user8612457

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:37:39