使用BeautifulSoup根据类名筛选爬取结果时返回空列表的问题
解决BeautifulSoup无法找到指定类名元素的问题
我来帮你分析下为什么代码返回空列表,以及怎么解决这个问题:
问题根源
你请求的这个页面的协会列表是通过JavaScript动态加载的。当你用requests.get()获取页面时,得到的只是静态的HTML框架,那些带有annuaire-item类的元素还没有被JS渲染出来,所以BeautifulSoup自然找不到它们。
你可以自己验证:把response.text保存成HTML文件打开,查看源码后会发现,里面根本没有annuaire-item相关的内容。
另外你代码里还有个小细节:循环变量用了class_,这和BeautifulSoup的参数名重名了,虽然不影响功能,但建议改成div这类变量名,避免混淆。
解决方案
这里提供两种可行的解决办法:
方法1:直接请求动态数据的API接口(推荐)
通过浏览器开发者工具的「网络」面板,能找到页面加载协会数据的真实API。直接请求这个接口可以拿到结构化的JSON数据,不需要解析HTML,效率更高:
import requests import pandas as pd # 页面加载协会数据的真实API地址 api_url = "https://www.ville-saintmandrier.fr/wp-json/wp/v2/associations?per_page=100" response = requests.get(api_url) if response.ok: data = response.json() names = [item["title"]["rendered"] for item in data] # 如需其他字段,可从item中提取,比如协会描述:item["content"]["rendered"] df = pd.DataFrame({"协会名称": names}) print(df)
方法2:用Selenium模拟浏览器加载页面
如果不想找API,可以用Selenium模拟浏览器打开页面,等待JS渲染完成后再抓取内容:
首先需要安装Selenium和对应浏览器驱动(比如ChromeDriver):
pip install selenium
然后修改代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import pandas as pd url = 'https://www.ville-saintmandrier.fr/acces-rapide/associations-mandreennes/#results' driver = webdriver.Chrome() # 用其他浏览器的话,替换成对应驱动 driver.get(url) # 等待目标元素加载完成,最多等待10秒 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "annuaire-item")) ) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') divs = soup.findAll(class_="annuaire-item") names = [] for div in divs: names.append(div.find("h3").text.strip()) print(div.text.strip()) df = pd.DataFrame({"协会名称": names}) print(df) finally: driver.quit()
小提示
优先用API的方法,因为网站的HTML结构可能会频繁改动,但API接口一般更稳定,维护成本更低。
内容的提问来源于stack exchange,提问作者Lilian Carion
相关产品推荐
相关产品推荐

