使用Selenium提取网页<h2>元素文本遇问题,求解决方案与优化
问题描述
我正在开发一个项目,需要批量处理邮政编码(后续会改为列表形式),将其输入CityFibre网站,循环处理每个条目并把网站返回结果保存到CSV文件。目前在提取网页中<h2>元素的文本内容时遇到问题,目标提取内容为:Great news! You can connect to the CityFibre network。以下是我编写的代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support import expected_conditions as EC import time driver = webdriver.Firefox() driver.get('https://cityfibre.com/') assert 'CityFibre' in driver.title time.sleep(2.5) cookies = WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.ID, "onetrust-accept-btn-handler"))) cookies.click() PostCode = driver.find_element(By.XPATH,("//input[@class='w-full border rounded-lg border-grey-950 bg-white text-grey-950 ac-outline focus:border-cf-blue placeholder-grey-500 md:text-base lg:rounded-xl lg:rounded-tl-none lg:rounded-tl-none rounded-tl-none pr-32 pl-11 py-5']")); PostCode.click() PostCode.send_keys('EH17 7RY') PostCode.send_keys(Keys.ENTER) time.sleep(2.5) PostCode.send_keys(Keys.TAB) PostCode.send_keys(Keys.ENTER) time.sleep(15) message = driver.find_element(By.TAG_NAME,('h2')[0]).text print(message) print('The End')
问题解决思路
- 修复
<h2>元素定位错误:原代码中(By.TAG_NAME,('h2')[0])是错误写法,('h2')[0]会取字符串'h2'的第一个字符'h',导致程序尝试定位<h>标签而非<h2>。正确写法应为By.TAG_NAME, 'h2';如果页面存在多个<h2>,需确认目标元素的索引(比如用find_elements(By.TAG_NAME, 'h2')[index]),或结合父容器、其他属性编写更精准的定位器(如//div[@class='result-container']/h2)。 - 补全缺失的导入:原代码使用了
WebDriverWait但未导入,需添加from selenium.webdriver.support.ui import WebDriverWait。 - 替换固定等待为显式等待:原代码大量使用
time.sleep(),这种固定等待方式不稳定(网络慢时会超时,网络快时浪费时间)。应改用WebDriverWait配合expected_conditions等待元素可见/可交互,确保页面加载完成后再执行操作。
代码优化建议
- 优化元素定位器:避免依赖冗长易变的class属性,改用更稳定的定位方式(如placeholder、name属性)。
- 封装逻辑为函数:方便后续批量处理邮政编码列表。
- 添加异常处理:防止单个邮政编码处理失败导致程序中断。
- 实现CSV保存功能:使用Python内置
csv模块将结果写入文件。
优化后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait import csv def check_postcode(driver, postcode): try: # 等待并点击接受cookie按钮 cookie_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "onetrust-accept-btn-handler")) ) cookie_btn.click() # 定位邮政编码输入框(使用placeholder更稳定) postcode_input = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//input[@placeholder='Enter your postcode']")) ) postcode_input.clear() postcode_input.send_keys(postcode) postcode_input.send_keys(Keys.ENTER) # 等待地址选择列表加载完成,再执行选择操作 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "address-list")) # 可根据实际页面元素调整 ) postcode_input.send_keys(Keys.TAB) postcode_input.send_keys(Keys.ENTER) # 等待结果h2元素可见并提取文本 result_h2 = WebDriverWait(driver, 15).until( EC.visibility_of_element_located((By.TAG_NAME, "h2")) ) return result_h2.text except Exception as e: print(f"处理邮政编码 {postcode} 时出错: {str(e)}") return "处理失败" if __name__ == "__main__": # 待处理的邮政编码列表 postcodes = ['EH17 7RY', 'WC1X 9NT', '其他测试邮编'] # 初始化浏览器 driver = webdriver.Firefox() driver.get('https://cityfibre.com/') assert 'CityFibre' in driver.title # 准备CSV文件并写入表头 with open('cityfibre_results.csv', 'w', newline='', encoding='utf-8') as csvfile: writer = csv.writer(csvfile) writer.writerow(['邮政编码', '查询结果']) # 循环处理每个邮政编码 for postcode in postcodes: result = check_postcode(driver, postcode) writer.writerow([postcode, result]) # 重置页面以便处理下一个邮编 driver.get('https://cityfibre.com/') driver.quit() print('处理完成,结果已保存至cityfibre_results.csv')
内容的提问来源于stack exchange,提问作者Lee Murray
相关产品推荐
相关产品推荐

