使用BeautifulSoup爬取homeshopping.pk手机价格遇返回空对象问题求助
解决方法:获取homeshopping.pk手机价格的正确姿势
首先,你遇到的find('div','ProductList')返回空的问题,大概率是两个原因:要么是页面元素的类名不对,要么是网站用JavaScript动态加载了产品内容,而requests只能拿到初始的静态HTML,根本没获取到产品数据。
我来一步步帮你解决:
1. 先排查静态页面的元素选择器
首先,打开目标网站的开发者工具(按F12),查看产品列表的实际HTML结构。实际的产品容器并不是ProductList这个类名,每个产品都在带有class="product-item"的div里,整个产品列表的父容器也有对应的标识。如果是静态内容的话,你可以修改选择器试试:
import requests import bs4 html_page = requests.get('https://homeshopping.pk/categories/Mobile-Phones-Price-Pakistan') html_page.raise_for_status() soup = bs4.BeautifulSoup(html_page.text, features='lxml') # 找到所有产品项 product_items = soup.find_all('div', class_='product-item') for item in product_items: # 获取产品名称 name = item.find('h3', class_='product-title').get_text(strip=True) # 获取价格 price = item.find('span', class_='price').get_text(strip=True) print(f"产品:{name},价格:{price}")
2. 如果静态方式不行,说明是动态加载(大概率是这个情况)
很多电商网站会用AJAX异步加载产品数据,这时候requests获取的初始HTML里根本没有产品内容,自然找不到元素。这时候可以用两种方法:
方法A:抓包找API接口
打开开发者工具的「网络」标签,刷新页面,过滤XHR请求,找返回产品数据的接口。直接请求这个接口就能拿到JSON格式的产品数据,比解析HTML方便多了。举个例子(实际需要你自己抓包确认接口地址):
import requests # 替换成你抓包找到的实际API接口 api_url = "https://homeshopping.pk/api/categories/Mobile-Phones-Price-Pakistan/products" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(api_url, headers=headers) response.raise_for_status() products = response.json() for product in products: print(f"产品:{product['name']},价格:{product['price']}")
方法B:用Selenium模拟浏览器加载
如果找不到API接口,或者接口有反爬限制,就用Selenium模拟真实浏览器打开页面,等待内容加载完成后再解析:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import bs4 # 初始化浏览器(需要下载对应浏览器的驱动,比如Chrome的chromedriver) driver = webdriver.Chrome() driver.get('https://homeshopping.pk/categories/Mobile-Phones-Price-Pakistan') # 等待产品列表加载完成,最多等10秒 wait = WebDriverWait(driver, 10) wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'product-item'))) # 获取页面源码 html = driver.page_source soup = bs4.BeautifulSoup(html, features='lxml') # 解析产品和价格 product_items = soup.find_all('div', class_='product-item') for item in product_items: name = item.find('h3', class_='product-title').get_text(strip=True) price = item.find('span', class_='price').get_text(strip=True) print(f"产品:{name},价格:{price}") # 关闭浏览器 driver.quit()
额外提示
- 记得添加
User-Agent请求头,避免被网站识别为爬虫而拦截。 - 如果网站有反爬机制,可能需要添加延时、使用代理等策略。
内容的提问来源于stack exchange,提问作者Hamza Ashes
相关产品推荐
相关产品推荐

