网页爬取故障求助:无法提取目标网站下拉列表内容
网页爬取问题排查与解决
问题背景
尝试爬取以下两个网站的公司名称下拉列表:
- https://secure.ethicspoint.com/domain/en/default_reporter.asp
- https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization
使用requests静态爬取和Selenium动态爬取均无法正常提取目标元素:第一个网站动态爬取报错,第二个网站爬取结果为国家列表而非预期的公司名称。
错误日志
DevTools listening on ws://127.0.0.1:53501/devtools/browser/028c6371-d9c3-4a13-83e1-2d7f598da093 Attempting static scraping for https://secure.ethicspoint.com/domain/en/default_reporter.asp... No company names found in static content. Static scraping failed, attempting dynamic scraping for https://secure.ethicspoint.com/domain/en/default_reporter.asp... Error in dynamic scraping: Message: no such element: Unable to locate element: {"method":"tag name","selector":"select"} (Session info: chrome=129.0.6668.101); For documentation on this error, please visit: https://www.selenium.dev/documentation/webdriver/troubleshooting/errors#no-such-element-exception Stacktrace: GetHandleVerifier [0x00AA5523+24195] (No symbol) [0x00A3AA04] (No symbol) [0x00932093] (No symbol) [0x00976ED2] (No symbol) [0x0097711B] (No symbol) [0x009B76F2] (No symbol) [0x0099AB84] (No symbol) [0x009B5280] (No symbol) [0x0099A8D6] (No symbol) [0x0096BA27] (No symbol) [0x0096C43D] GetHandleVerifier [0x00D6CE13+2938739] GetHandleVerifier [0x00DBEC69+3274185] GetHandleVerifier [0x00B309C2+594722] GetHandleVerifier [0x00B37EDC+624700] (No symbol) [0x00A437CD] (No symbol) [0x00A40528] (No symbol) [0x00A406C5] (No symbol) [0x00A32CA6] BaseThreadInitThunk [0x7648FCC9+25] RtlGetAppContainerNamedObjectPath [0x779C80CE+286] RtlGetAppContainerNamedObjectPath [0x779C809E+238] No companies found on https://secure.ethicspoint.com/domain/en/default_reporter.asp. Attempting static scraping for https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization... Companies found on https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization: - Select your location - Albania - Andorra - Angola - Antigua and Barbuda - Argentina
爬取代码
import requests from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By import time driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) def static_scrape(url): try: response = requests.get(url) if response.status_code == 200: soup = BeautifulSoup(response.text, 'html.parser') options = soup.find_all('option') if options: companies = [option.text.strip() for option in options if option.text.strip()] return companies else: print("No company names found in static content.") return None else: print(f"Failed to retrieve webpage. Status code: {response.status_code}") return None except Exception as e: print(f"Error in static scraping: {e}") return None def dynamic_scrape(url): try: driver.get(url) time.sleep(5) dropdown = driver.find_element(By.TAG_NAME, 'select') options = dropdown.find_elements(By.TAG_NAME, 'option') companies = [option.text.strip() for option in options if option.text.strip()] return companies except Exception as e: print(f"Error in dynamic scraping: {e}") return None def scrape_companies(url): print(f"Attempting static scraping for {url}...") companies = static_scrape(url) if companies is None: print(f"Static scraping failed, attempting dynamic scraping for {url}...") companies = dynamic_scrape(url) return companies urls = [ 'https://secure.ethicspoint.com/domain/en/default_reporter.asp', 'https://app.convercent.com/en-us/Anonymous/IssueIntake/IdentifyOrganization' ] for url in urls: companies = scrape_companies(url) if companies: print(f"\nCompanies found on {url}:") for company in companies: print(f"- {company}") else: print(f"No companies found on {url}.\n") driver.quit()
问题分析与解决
1. 第一个网站(ethicspoint.com)问题解决
问题原因
- 静态爬取:下拉列表为动态加载生成,
requests获取的静态HTML中无select元素及选项内容。 - 动态爬取:该网站未使用原生
<select>标签,而是用自定义UI组件模拟下拉,原代码通过标签名定位必然失败;且固定等待时间可能不足以完成页面渲染。
解决步骤
- 替换固定等待为显式等待,确保元素加载完成。
- 通过CSS选择器定位自定义下拉触发按钮,点击后抓取选项列表。
修改后的动态爬取函数:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def dynamic_scrape_ethicspoint(url): try: driver.get(url) # 等待下拉触发按钮加载并点击 dropdown_trigger = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div.select2-selection--single")) ) dropdown_trigger.click() # 等待选项列表加载并提取内容 options = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, "ul.select2-results__options li")) ) companies = [option.text.strip() for option in options if option.text.strip()] return companies except Exception as e: print(f"Error in dynamic scraping: {e}") return None
2. 第二个网站(convercent.com)问题解决
问题原因
原代码抓取的是地区选择下拉列表,该网站需先选择地区,才会加载对应地区的公司列表。
解决步骤
- 先定位地区下拉列表,选择目标地区。
- 等待公司列表加载完成后,再抓取公司名称选项。
修改后的动态爬取函数:
from selenium.webdriver.support.ui import Select def dynamic_scrape_convercent(url): try: driver.get(url) # 等待地区下拉加载并选择目标地区(示例为美国) location_dropdown = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "LocationId")) ) select_location = Select(location_dropdown) select_location.select_by_visible_text("United States") # 等待公司列表加载 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "OrganizationId")) ) # 提取公司选项 company_dropdown = driver.find_element(By.ID, "OrganizationId") options = company_dropdown.find_elements(By.TAG_NAME, "option") companies = [option.text.strip() for option in options if option.text.strip() and not option.text.startswith("Select")] return companies except Exception as e: print(f"Error in dynamic scraping: {e}") return None
通用优化建议
- 优先使用显式等待替代固定
time.sleep,提升爬取稳定性与效率。 - 针对非原生下拉组件,需分析页面DOM结构,定位实际触发元素与选项容器。
- 检查页面是否包含
iframe,若目标元素在iframe内,需先切换到iframe再进行元素定位。
内容的提问来源于stack exchange,提问作者king
相关产品推荐
相关产品推荐

