使用BeautifulSoup抓取店铺目录:名称提取错误,描述正确
问题分析
你当前的代码错误在于,soup.find_all(class_="table table-hover")获取的是整个店铺列表的表格,而非单个店铺的行数据。循环表格时提取的td[0:2]是表格的前两个单元格,并非对应每个店铺的名称,导致名称提取重复且错误。
修正后的代码
import requests from bs4 import BeautifulSoup url = "https://www.jurongpoint.com.sg/store-directory/?level=&cate=Food+%26+Beverage&page=1" data = requests.get(url) soup = BeautifulSoup(data.content, "html.parser") # 获取店铺列表的表格 shop_table = soup.find(class_="table table-hover") # 获取表格中所有店铺行(跳过表头行) shop_rows = shop_table.find_all("tr")[1:] for row in shop_rows: # 提取店铺名称:第一个td中的a标签文本 shop_name = row.find("td").find("a").get_text(strip=True) # 提取店铺描述:该行中的col-9 div文本 shop_desc = row.find("div", class_="col-9").get_text(strip=True) print(f"店铺名称: {shop_name}") print(f"店铺描述: {shop_desc}") print("-" * 50)
代码说明
- 先定位到唯一的店铺表格,再获取表格内的所有行,跳过第一行表头。
- 对每一行(单个店铺),从第一个
<td>里的<a>标签提取店铺名称,确保精准对应。 - 使用
get_text(strip=True)去除文本中的多余空格和换行,让输出更整洁。
内容的提问来源于stack exchange,提问作者Bee
相关产品推荐
相关产品推荐

