You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium和BeautifulSoup无法抓取Craigslist页面指定元素的问题

问题

尝试抓取Craigslist巴尔的摩站的搜索结果页面,开发者工具里能看到class为cl-search-result的li标签(共42条结果),但用Selenium和BeautifulSoup抓取时,返回的soup里找不到这些元素,页面源码和开发者工具显示的HTML完全不一致。

原代码如下:

import time
import datetime
from collections import namedtuple
import selenium.webdriver as webdriver
from selenium.webdriver.firefox.service import Service
from selenium.webdriver.support.ui import Select
from selenium.webdriver.firefox.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.common.exceptions import ElementNotInteractableException
from bs4 import BeautifulSoup
import pandas as pd
import os

user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/109.0'
firefox_driver_path = os.path.join(os.getcwd(), 'geckodriver.exe')
firefox_service = Service(firefox_driver_path)
firefox_option = Options()
firefox_option.set_preference('general.useragent.override', user_agent)
browser = webdriver.Firefox(service=firefox_service, options=firefox_option)
browser.implicitly_wait(7)


url = 'https://baltimore.craigslist.org/search/sss#search=1~list~0~0'
browser.get(url)

soup = BeautifulSoup(browser.page_source, 'html.parser') 
print(soup)
posts_html= soup.find_all('li', {'class': 'cl-search-result'})
   
print('Collected {0} listings'.format(len(posts_html)))

解决方案

1. 问题根源

Craigslist的搜索结果是动态渲染的,原代码里的隐式等待(implicitly_wait)只是等待元素可交互,没法覆盖动态内容加载的延迟;另外URL里的锚点参数#search=1~list~0~0可能触发特殊加载逻辑,导致页面没正常渲染出结果。

2. 具体修复步骤

  • 换成显式等待:精准等待目标结果元素加载完成,比隐式等待更可靠
  • 简化URL:去掉锚点参数,用基础搜索URL避免加载异常
  • 可选添加页面滚动:部分动态内容需要滚动才会加载完全
  • 处理可能的反爬验证:如果遇到弹窗验证,需手动或代码触发点击

3. 修改后的代码示例

import time
import datetime
from collections import namedtuple
import selenium.webdriver as webdriver
from selenium.webdriver.firefox.service import Service
from selenium.webdriver.support.ui import Select
from selenium.webdriver.firefox.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.common.keys import Keys
from selenium.common.exceptions import ElementNotInteractableException, TimeoutException
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd
import os

user_agent = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/109.0'
firefox_driver_path = os.path.join(os.getcwd(), 'geckodriver.exe')
firefox_service = Service(firefox_driver_path)
firefox_option = Options()
firefox_option.set_preference('general.useragent.override', user_agent)
# 可选:禁用图片加载,提升爬取速度
firefox_option.set_preference('permissions.default.image', 2)
browser = webdriver.Firefox(service=firefox_service, options=firefox_option)

try:
    # 用基础搜索URL,去掉锚点参数
    url = 'https://baltimore.craigslist.org/search/sss'
    browser.get(url)

    # 显式等待搜索结果元素出现,最长等15秒
    wait = WebDriverWait(browser, 15)
    wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'cl-search-result')))

    # 滚动页面到底部,确保所有结果加载完成
    browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(2)  # 给滚动后的加载留缓冲时间

    # 获取完全渲染后的页面源码
    soup = BeautifulSoup(browser.page_source, 'html.parser')
    posts_html = soup.find_all('li', {'class': 'cl-search-result'})

    print(f'Collected {len(posts_html)} listings')
finally:
    # 无论是否成功,都关闭浏览器
    browser.quit()

4. 额外排查点

  • 检查浏览器是否在无头模式下异常:如果用无头模式,确保Firefox版本和geckodriver兼容
  • 验证user-agent是否生效:在soup里查找<meta name="user-agent">标签,确认和设置的一致
  • 如果仍无法获取结果,可能遇到了Cloudflare反爬:这种情况可以尝试用undetected-chromedriver替代普通Selenium

内容的提问来源于stack exchange,提问作者Tendekai Muchenje

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 04:35:46