You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取HTML中select标签的有效option值与文本?

问题描述

需要抓取目标页面中<select id="aws">下的<option>标签的value值和站点全名,分别存入两个独立列表,同时过滤掉value为空或站点名无效的选项(比如分隔符、空内容的选项)。

目标HTML片段(仅展示前几项):

<select title="Please select Automatic Weather Observations" id="aws" onchange="changeYearDropDown2()">
    <option value="">===Manned Weather Station===</option>
    <option value="HKO">Hong Kong Observatory</option>
    <option value="HKA">Hong Kong International Airport</option>
    <option value=""></option>
    <option value="">===Automatic Weather Station===</option>
    <option value="BR1">Beas River</option>
    <option value="BHD">Bluff Head</option>
    <option value="CP1">Central Pier</option>
    <option value="CCH">Cheung Chau</option>
    <option value="CPH">Ching Pak House(Tsing Yi)</option>
</select>

尝试过的方法均失败:

使用BeautifulSoup

import requests
from bs4 import BeautifulSoup as bs
  
url = 'https://www.hko.gov.hk/en/cis/climat.htm'
  
req = requests.get(url)
soup = bs(req.text, 'html.parser')

# method 1
print(soup.select)    # 输出整个文档
print(soup.option)    # 空列表

# method 2
titles = soup.find_all('option')
print(titles)    # 空列表

使用Selenium

from selenium.webdriver.chrome.service import Service
from selenium import webdriver
from selenium.webdriver.common.by import By

service = Service(executable_path="path/chromedriver_win32/chromedriver.exe")
# 初始化浏览器驱动
with webdriver.Chrome(service=service) as driver:
    driver.get(url)

# method 3: CSS选择器定位
    myDiv = driver.find_element(By.CSS_SELECTOR, "select option")    # find_elements报错
    print(myDiv.get_attribute("outerHTML"))
    print(myDiv.get_attribute("innerHTML")) # 输出<select></select>

# method 4: 标签名定位
    myDiv2 = driver.find_element(By.TAG_NAME, 'select')
    print(myDiv2.get_attribute("outerHTML"))    
    # 能返回包含站点信息的长HTML字符串,但不好处理

# method 5: XPATH循环
    list = []
    for r in range(1,74):    # 共73个选项
        value = driver.find_element(By.XPATH, "/html/body/div[2]/div[2]/div/div/div[4]/div/div[3]/div[1]/div/div[2]/div[2]/div[1]/div/ul/li[1]/select/option["+str(r)+"]").text
        list.append(value)
    print(list)
    # 只能获取文本,且返回空列表
解决方案

原因分析

  1. BeautifulSoup失败原因:目标页面的<select id="aws">是通过JavaScript动态加载的,直接用requests.get()只能获取初始HTML,无法拿到JS渲染后的内容。
  2. Selenium失败原因:没有等待页面加载完成就直接定位元素,导致元素还未渲染;XPATH路径太固定,页面结构变化就会失效。

方法一:使用Selenium + 显式等待

利用Selenium的显式等待,等目标元素加载完成后再抓取,同时过滤无效选项:

from selenium.webdriver.chrome.service import Service
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

service = Service(executable_path="path/chromedriver_win32/chromedriver.exe")
value_list = []
name_list = []

with webdriver.Chrome(service=service) as driver:
    driver.get('https://www.hko.gov.hk/en/cis/climat.htm')
    # 显式等待select元素加载完成,最多等10秒
    select_element = WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.ID, "aws"))
    )
    # 获取所有option元素
    options = select_element.find_elements(By.TAG_NAME, "option")
    
    # 遍历过滤有效选项
    for opt in options:
        val = opt.get_attribute("value")
        name = opt.text.strip()
        # 过滤条件:value不为空,且名称不是分隔符、不是空字符串
        if val and name and "===" not in name:
            value_list.append(val)
            name_list.append(name)

print("Value列表:", value_list)
print("站点名称列表:", name_list)

方法二:直接调用页面API(更高效)

观察页面网络请求,发现站点数据是通过API接口返回的,直接请求API可避免JS渲染问题:

import requests

# 站点数据API接口
api_url = "https://www.hko.gov.hk/cis/awsJsonEN.json"
response = requests.get(api_url)
data = response.json()

value_list = []
name_list = []

# 解析返回的JSON数据
for item in data:
    # 提取value和名称,"code"对应option的value,"name"对应站点全名
    value_list.append(item["code"])
    name_list.append(item["name"])

print("Value列表:", value_list)
print("站点名称列表:", name_list)

该方法无需加载整个页面,直接获取结构化数据,效率远高于页面渲染抓取。

内容的提问来源于stack exchange,提问作者abcfhy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 20:24:53