使用Python爬取SIPSA展商网站返回空数据问题求助
问题排查与解决方法
1. 核心问题:选择器不匹配页面元素
你的代码无法抓取数据,大概率是因为使用的CSS选择器和当前网站的HTML结构不匹配,导致soup.select(".exposant-list-item")返回空列表,没有可遍历的展商元素。
快速调试验证
在获取soup后添加以下代码,确认选择器是否有效:
# 检查是否找到展商容器 exhibitors = soup.select(".exposant-list-item") print(f"当前页面找到{len(exhibitors)}个展商") if not exhibitors: # 打印页面片段,确认是否获取到正确内容 print(response.text[:500])
如果输出当前页面找到0个展商,需要手动检查网页结构:
- 打开目标页面,右键选择「检查」,定位到展商卡片元素,查看它的实际class属性(可能已变更为
exhibitor-card或其他名称)。 - 依次核对展商名称、行业、产品等子元素的class,比如
.title是否还对应展商名称,.secteur-activite是否存在。
2. 处理动态加载页面
如果静态HTML中没有展商数据,说明页面是通过JavaScript动态渲染的,requests只能获取到初始空页面,无法抓取渲染后的内容。此时需改用selenium模拟浏览器加载:
修改后的示例代码
import pandas as pd from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time from bs4 import BeautifulSoup data = [] max_pages = 30 # 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver) driver = webdriver.Chrome() for page in range(1, max_pages + 1): print("正在爬取第", page, "页") url = f"https://www.sipsa-filaha.com/fr/exposant/?page={page}" driver.get(url) try: # 等待展商列表加载完成(超时10秒) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "exposant-list-item")) ) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, "html.parser") for exhibitor in soup.select(".exposant-list-item"): # 增加空值判断,避免因元素缺失报错 name = exhibitor.select_one(".title").get_text(strip=True) if exhibitor.select_one(".title") else "无数据" sector = exhibitor.select_one(".secteur-activite").get_text(strip=True) if exhibitor.select_one(".secteur-activite") else "无数据" services_products = exhibitor.select_one(".services-produits").get_text(strip=True) if exhibitor.select_one(".services-produits") else "无数据" country = exhibitor.select_one(".country").get_text(strip=True) if exhibitor.select_one(".country") else "无数据" contacts = exhibitor.select_one(".contact").get_text(strip=True) if exhibitor.select_one(".contact") else "无数据" data.append((name, sector, services_products, country, contacts)) time.sleep(2) except Exception as e: print("爬取出错:", e) break driver.quit() # 保存数据 df = pd.DataFrame(data, columns=["Name", "Sector", "Services/Products", "Country", "Contacts"]) excel_file_path = r"C:\Users\LENOVO\Desktop\scraping_sante\sipsa_exhibitors_data.xlsx" df.to_excel(excel_file_path, index=False) print("数据已保存至:", excel_file_path)
3. 优化请求头避免被拦截
部分网站会通过请求头识别爬虫,可补充更完整的请求头字段:
headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", "Accept-Language": "fr-FR,fr;q=0.8,en-US;q=0.5,en;q=0.3", "Referer": "https://www.sipsa-filaha.com/fr/" }
4. 确认页码范围有效性
检查你设置的max_pages=30是否超过网站实际的展商页数,比如手动访问?page=30,确认页面是否真的有展商数据。
内容的提问来源于stack exchange,提问作者Wiam.07 Lazazi
相关产品推荐
相关产品推荐

