Selenium抓取下拉菜单数据如何将结果存储为列表格式
Selenium提取下拉菜单选项为列表的实现方法
你直接读取select元素的text属性时,会把元素内部所有文本(包含HTML缩进产生的换行、多余空格)拼接为整段字符串,无法直接得到结构化结果,可通过以下两种方式实现需求:
方法1:遍历option子元素提取
定位到select元素后,查找其下所有<option>子节点,逐个提取文本并去除首尾空白字符即可:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 定位下拉框父元素 select_el = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "filter")) ) # 获取所有option子元素 all_options = select_el.find_elements(By.TAG_NAME, "option") # 清洗文本得到结构化列表 option_list = [item.text.strip() for item in all_options if item.text.strip()]
运行后option_list的输出结果为:
['last 6 months', '2022', '2021', '2020']
如果只需要提取年份选项,可以通过option的value属性做过滤:
year_list = [item.text.strip() for item in all_options if item.get_attribute("value").startswith("year-")]
最终得到纯年份列表:['2022', '2021', '2020']
方法2:使用Selenium内置Select类
Selenium提供了专门处理下拉选择框的Select工具类,封装了下拉框常用操作,不需要手动查找option元素:
from selenium.webdriver.support.ui import Select select_el = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "filter")) ) # 初始化下拉框对象 dropdown = Select(select_el) # 直接获取所有选项并清洗文本 option_list = [item.text.strip() for item in dropdown.options]
Select类还支持按索引、按value值、按可见文本直接选中对应选项,处理标准select下拉框时稳定性更高。
内容的提问来源于stack exchange,提问作者Pisa Ponente
相关产品推荐
相关产品推荐

