You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取故障求助:无法提取目标网站下拉列表内容

网页爬取问题排查与解决

问题背景

尝试爬取以下两个网站的公司名称下拉列表:

  • https://secure.ethicspoint.com/domain/en/default_reporter.asp
  • https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization

使用requests静态爬取和Selenium动态爬取均无法正常提取目标元素:第一个网站动态爬取报错,第二个网站爬取结果为国家列表而非预期的公司名称。

错误日志

DevTools listening on ws://127.0.0.1:53501/devtools/browser/028c6371-d9c3-4a13-83e1-2d7f598da093
Attempting static scraping for https://secure.ethicspoint.com/domain/en/default_reporter.asp...
No company names found in static content.
Static scraping failed, attempting dynamic scraping for https://secure.ethicspoint.com/domain/en/default_reporter.asp...
Error in dynamic scraping: Message: no such element: Unable to locate element: {"method":"tag name","selector":"select"}
  (Session info: chrome=129.0.6668.101); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception
Stacktrace:
        GetHandleVerifier [0x00AA5523+24195]
        (No symbol) [0x00A3AA04]
        (No symbol) [0x00932093]
        (No symbol) [0x00976ED2]
        (No symbol) [0x0097711B]
        (No symbol) [0x009B76F2]
        (No symbol) [0x0099AB84]
        (No symbol) [0x009B5280]
        (No symbol) [0x0099A8D6]
        (No symbol) [0x0096BA27]
        (No symbol) [0x0096C43D]
        GetHandleVerifier [0x00D6CE13+2938739]
        GetHandleVerifier [0x00DBEC69+3274185]
        GetHandleVerifier [0x00B309C2+594722]
        GetHandleVerifier [0x00B37EDC+624700]
        (No symbol) [0x00A437CD]
        (No symbol) [0x00A40528]
        (No symbol) [0x00A406C5]
        (No symbol) [0x00A32CA6]
        BaseThreadInitThunk [0x7648FCC9+25]
        RtlGetAppContainerNamedObjectPath [0x779C80CE+286]
        RtlGetAppContainerNamedObjectPath [0x779C809E+238]

No companies found on https://secure.ethicspoint.com/domain/en/default_reporter.asp.

Attempting static scraping for https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization...

Companies found on https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization:
- Select your location
- Albania
- Andorra
- Angola
- Antigua and Barbuda
- Argentina

爬取代码

import requests
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.chrome.service import Service as ChromeService
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
import time

driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install()))

def static_scrape(url):
   
    try:
       
        response = requests.get(url)

        if response.status_code == 200:
            
            soup = BeautifulSoup(response.text, 'html.parser')

            options = soup.find_all('option')

            if options:
                companies = [option.text.strip() for option in options if option.text.strip()]
                return companies
            else:
                print("No company names found in static content.")
                return None
        else:
            print(f"Failed to retrieve webpage. Status code: {response.status_code}")
            return None
    except Exception as e:
        print(f"Error in static scraping: {e}")
        return None

def dynamic_scrape(url):
   
    try:
        
        driver.get(url)

        
        time.sleep(5)

        dropdown = driver.find_element(By.TAG_NAME, 'select')

        
        options = dropdown.find_elements(By.TAG_NAME, 'option')

       
        companies = [option.text.strip() for option in options if option.text.strip()]

        return companies

    except Exception as e:
        print(f"Error in dynamic scraping: {e}")
        return None

def scrape_companies(url):
    
    print(f"Attempting static scraping for {url}...")
    companies = static_scrape(url)

    if companies is None:
        print(f"Static scraping failed, attempting dynamic scraping for {url}...")
        companies = dynamic_scrape(url)

    return companies


urls = [
    'https://secure.ethicspoint.com/domain/en/default_reporter.asp',
    'https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization'
]


for url in urls:
    companies = scrape_companies(url)

    if companies:
        print(f"\nCompanies found on {url}:")
        for company in companies:
            print(f"- {company}")
    else:
        print(f"No companies found on {url}.\n")


driver.quit()

问题分析与解决

1. 第一个网站(ethicspoint.com)问题解决

问题原因

  • 静态爬取:下拉列表为动态加载生成,requests获取的静态HTML中无select元素及选项内容。
  • 动态爬取:该网站未使用原生<select>标签,而是用自定义UI组件模拟下拉,原代码通过标签名定位必然失败;且固定等待时间可能不足以完成页面渲染。

解决步骤

  • 替换固定等待为显式等待,确保元素加载完成。
  • 通过CSS选择器定位自定义下拉触发按钮,点击后抓取选项列表。

修改后的动态爬取函数:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def dynamic_scrape_ethicspoint(url):
    try:
        driver.get(url)
        # 等待下拉触发按钮加载并点击
        dropdown_trigger = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "div.select2-selection--single"))
        )
        dropdown_trigger.click()
        # 等待选项列表加载并提取内容
        options = WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CSS_SELECTOR, "ul.select2-results__options li"))
        )
        companies = [option.text.strip() for option in options if option.text.strip()]
        return companies
    except Exception as e:
        print(f"Error in dynamic scraping: {e}")
        return None

2. 第二个网站(convercent.com)问题解决

问题原因

原代码抓取的是地区选择下拉列表,该网站需先选择地区,才会加载对应地区的公司列表。

解决步骤

  • 先定位地区下拉列表,选择目标地区。
  • 等待公司列表加载完成后,再抓取公司名称选项。

修改后的动态爬取函数:

from selenium.webdriver.support.ui import Select

def dynamic_scrape_convercent(url):
    try:
        driver.get(url)
        # 等待地区下拉加载并选择目标地区(示例为美国)
        location_dropdown = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "LocationId"))
        )
        select_location = Select(location_dropdown)
        select_location.select_by_visible_text("United States")
        # 等待公司列表加载
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "OrganizationId"))
        )
        # 提取公司选项
        company_dropdown = driver.find_element(By.ID, "OrganizationId")
        options = company_dropdown.find_elements(By.TAG_NAME, "option")
        companies = [option.text.strip() for option in options if option.text.strip() and not option.text.startswith("Select")]
        return companies
    except Exception as e:
        print(f"Error in dynamic scraping: {e}")
        return None

通用优化建议

  • 优先使用显式等待替代固定time.sleep,提升爬取稳定性与效率。
  • 针对非原生下拉组件,需分析页面DOM结构,定位实际触发元素与选项容器。
  • 检查页面是否包含iframe,若目标元素在iframe内,需先切换到iframe再进行元素定位。

内容的提问来源于stack exchange,提问作者king

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 06:07:02