使用Beautiful Soup抓取黄页网站电话与网址无结果,求问题排查
问题分析与修复方案
问题根源
- 反爬拦截:未添加请求头模拟浏览器,网站返回的内容不包含目标商家数据。
- 选择器匹配错误:原代码使用的HTML元素class与当前页面实际结构不符,无法定位到电话和网址元素。
修复后的代码
import requests from bs4 import BeautifulSoup url = 'https://www.yellowpages.ca/search/si/1/coffee/Toronto+ON' # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = requests.get(url, headers=headers) # 校验请求是否成功 if response.status_code != 200: print(f"请求失败,状态码:{response.status_code}") else: soup = BeautifulSoup(response.text, 'html.parser') # 遍历商家列表项(匹配当前页面的class) for listing in soup.find_all('div', class_='listing_right_section'): # 提取电话号码(存储在data-phone属性中) phone_elem = listing.find('a', class_='mlr__item mlr__item--phone') if phone_elem: print(f"电话号码:{phone_elem.get('data-phone')}") # 提取商家网址(存储在href属性中) website_elem = listing.find('a', class_='mlr__item mlr__item--website') if website_elem: print(f"商家网址:{website_elem.get('href')}") print("--- 分割线 ---")
关键说明
- 请求头设置:添加
User-Agent让服务器认为请求来自常规浏览器,绕过基础反爬机制。 - 元素定位修正:当前页面的商家列表使用
listing_right_section类,电话和网址元素均为<a>标签,数据分别存储在data-phone和href属性中(而非直接文本),原代码的选择器和取值方式均不符合当前页面结构。 - 状态码校验:快速排查请求失败的情况(如IP被封禁、网络问题等)。
内容的提问来源于stack exchange,提问作者Sara
相关产品推荐
相关产品推荐

