You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取SIPSA展商网站返回空数据问题求助

问题排查与解决方法

1. 核心问题:选择器不匹配页面元素

你的代码无法抓取数据,大概率是因为使用的CSS选择器和当前网站的HTML结构不匹配,导致soup.select(".exposant-list-item")返回空列表,没有可遍历的展商元素。

快速调试验证

在获取soup后添加以下代码,确认选择器是否有效:

# 检查是否找到展商容器
exhibitors = soup.select(".exposant-list-item")
print(f"当前页面找到{len(exhibitors)}个展商")
if not exhibitors:
    # 打印页面片段,确认是否获取到正确内容
    print(response.text[:500])

如果输出当前页面找到0个展商,需要手动检查网页结构:

  • 打开目标页面,右键选择「检查」,定位到展商卡片元素,查看它的实际class属性(可能已变更为exhibitor-card或其他名称)。
  • 依次核对展商名称、行业、产品等子元素的class,比如.title是否还对应展商名称,.secteur-activite是否存在。

2. 处理动态加载页面

如果静态HTML中没有展商数据,说明页面是通过JavaScript动态渲染的,requests只能获取到初始空页面,无法抓取渲染后的内容。此时需改用selenium模拟浏览器加载:

修改后的示例代码

import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
from bs4 import BeautifulSoup

data = []
max_pages = 30

# 初始化Chrome浏览器(需提前下载对应版本的ChromeDriver)
driver = webdriver.Chrome()

for page in range(1, max_pages + 1):
    print("正在爬取第", page, "页")
    url = f"https://www.sipsa-filaha.com/fr/exposant/?page={page}"
    driver.get(url)
    
    try:
        # 等待展商列表加载完成(超时10秒)
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, "exposant-list-item"))
        )
        
        # 获取渲染后的页面源码
        soup = BeautifulSoup(driver.page_source, "html.parser")
        
        for exhibitor in soup.select(".exposant-list-item"):
            # 增加空值判断,避免因元素缺失报错
            name = exhibitor.select_one(".title").get_text(strip=True) if exhibitor.select_one(".title") else "无数据"
            sector = exhibitor.select_one(".secteur-activite").get_text(strip=True) if exhibitor.select_one(".secteur-activite") else "无数据"
            services_products = exhibitor.select_one(".services-produits").get_text(strip=True) if exhibitor.select_one(".services-produits") else "无数据"
            country = exhibitor.select_one(".country").get_text(strip=True) if exhibitor.select_one(".country") else "无数据"
            contacts = exhibitor.select_one(".contact").get_text(strip=True) if exhibitor.select_one(".contact") else "无数据"
            
            data.append((name, sector, services_products, country, contacts))
        
        time.sleep(2)
        
    except Exception as e:
        print("爬取出错:", e)
        break

driver.quit()

# 保存数据
df = pd.DataFrame(data, columns=["Name", "Sector", "Services/Products", "Country", "Contacts"])
excel_file_path = r"C:\Users\LENOVO\Desktop\scraping_sante\sipsa_exhibitors_data.xlsx"
df.to_excel(excel_file_path, index=False)
print("数据已保存至:", excel_file_path)

3. 优化请求头避免被拦截

部分网站会通过请求头识别爬虫,可补充更完整的请求头字段:

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
    "Accept-Language": "fr-FR,fr;q=0.8,en-US;q=0.5,en;q=0.3",
    "Referer": "https://www.sipsa-filaha.com/fr/"
}

4. 确认页码范围有效性

检查你设置的max_pages=30是否超过网站实际的展商页数,比如手动访问?page=30,确认页面是否真的有展商数据。

内容的提问来源于stack exchange,提问作者Wiam.07 Lazazi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 23:15:37