如何用Python提取HTML中select标签的有效option值与文本?
问题描述
需要抓取目标页面中<select id="aws">下的<option>标签的value值和站点全名,分别存入两个独立列表,同时过滤掉value为空或站点名无效的选项(比如分隔符、空内容的选项)。
目标HTML片段(仅展示前几项):
<select title="Please select Automatic Weather Observations" id="aws" onchange="changeYearDropDown2()"> <option value="">===Manned Weather Station===</option> <option value="HKO">Hong Kong Observatory</option> <option value="HKA">Hong Kong International Airport</option> <option value=""></option> <option value="">===Automatic Weather Station===</option> <option value="BR1">Beas River</option> <option value="BHD">Bluff Head</option> <option value="CP1">Central Pier</option> <option value="CCH">Cheung Chau</option> <option value="CPH">Ching Pak House(Tsing Yi)</option> </select>
尝试过的方法均失败:
使用BeautifulSoup
import requests from bs4 import BeautifulSoup as bs url = 'https://www.hko.gov.hk/en/cis/climat.htm' req = requests.get(url) soup = bs(req.text, 'html.parser') # method 1 print(soup.select) # 输出整个文档 print(soup.option) # 空列表 # method 2 titles = soup.find_all('option') print(titles) # 空列表
使用Selenium
from selenium.webdriver.chrome.service import Service from selenium import webdriver from selenium.webdriver.common.by import By service = Service(executable_path="path/chromedriver_win32/chromedriver.exe") # 初始化浏览器驱动 with webdriver.Chrome(service=service) as driver: driver.get(url) # method 3: CSS选择器定位 myDiv = driver.find_element(By.CSS_SELECTOR, "select option") # find_elements报错 print(myDiv.get_attribute("outerHTML")) print(myDiv.get_attribute("innerHTML")) # 输出<select></select> # method 4: 标签名定位 myDiv2 = driver.find_element(By.TAG_NAME, 'select') print(myDiv2.get_attribute("outerHTML")) # 能返回包含站点信息的长HTML字符串,但不好处理 # method 5: XPATH循环 list = [] for r in range(1,74): # 共73个选项 value = driver.find_element(By.XPATH, "/html/body/div[2]/div[2]/div/div/div[4]/div/div[3]/div[1]/div/div[2]/div[2]/div[1]/div/ul/li[1]/select/option["+str(r)+"]").text list.append(value) print(list) # 只能获取文本,且返回空列表
解决方案
原因分析
- BeautifulSoup失败原因:目标页面的
<select id="aws">是通过JavaScript动态加载的,直接用requests.get()只能获取初始HTML,无法拿到JS渲染后的内容。 - Selenium失败原因:没有等待页面加载完成就直接定位元素,导致元素还未渲染;XPATH路径太固定,页面结构变化就会失效。
方法一:使用Selenium + 显式等待
利用Selenium的显式等待,等目标元素加载完成后再抓取,同时过滤无效选项:
from selenium.webdriver.chrome.service import Service from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC service = Service(executable_path="path/chromedriver_win32/chromedriver.exe") value_list = [] name_list = [] with webdriver.Chrome(service=service) as driver: driver.get('https://www.hko.gov.hk/en/cis/climat.htm') # 显式等待select元素加载完成,最多等10秒 select_element = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "aws")) ) # 获取所有option元素 options = select_element.find_elements(By.TAG_NAME, "option") # 遍历过滤有效选项 for opt in options: val = opt.get_attribute("value") name = opt.text.strip() # 过滤条件:value不为空,且名称不是分隔符、不是空字符串 if val and name and "===" not in name: value_list.append(val) name_list.append(name) print("Value列表:", value_list) print("站点名称列表:", name_list)
方法二:直接调用页面API(更高效)
观察页面网络请求,发现站点数据是通过API接口返回的,直接请求API可避免JS渲染问题:
import requests # 站点数据API接口 api_url = "https://www.hko.gov.hk/cis/awsJsonEN.json" response = requests.get(api_url) data = response.json() value_list = [] name_list = [] # 解析返回的JSON数据 for item in data: # 提取value和名称,"code"对应option的value,"name"对应站点全名 value_list.append(item["code"]) name_list.append(item["name"]) print("Value列表:", value_list) print("站点名称列表:", name_list)
该方法无需加载整个页面,直接获取结构化数据,效率远高于页面渲染抓取。
内容的提问来源于stack exchange,提问作者abcfhy
相关产品推荐
相关产品推荐

