使用Selenium爬取PFF.com四分卫评分时文本获取失败问题
问题:Selenium爬取PFF四分卫评分时.text无法捕获内容,结果列表为空
使用Selenium爬取PFF.com的NFL 2022赛季常规赛四分卫评分,目标是获取所有四分卫的姓名及对应评分,但遇到问题:调用.text方法无法捕获目标元素的文本内容,且未触发NoSuchElementException异常,最终输出的评分列表和姓名列表均为空。
原代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from time import sleep service = Service(executable_path=r"C:\chromedriver.exe") op = webdriver.ChromeOptions() driver = webdriver.Chrome(service=service, options=op) driver.get("https://premium.pff.com/nfl/positions/2022/REG/passing?position=QB") sleep(2) sign_in = driver.find_element(By.XPATH, '/html/body/div/div/header/div[3]/button') sign_in.click() sleep(2) email = driver.find_element(By.XPATH, '/html/body/div/div/div/div/div/div/form/div[1]/input') email.send_keys(my_email) password = driver.find_element(By.XPATH, '/html/body/div/div/div/div/div/div/form/div[2]/input') password.send_keys(my_password) sleep(2) sign_in_2 = driver.find_element(By.XPATH, '/html/body/div/div/div/div/div/div/form/button') sign_in_2.click() sleep(2) all_off_grades = driver.find_elements(By.CSS_SELECTOR, '.kyber-table .kyber-grade-badge__info-text div') all_qb_names = driver.find_elements(By.CSS_SELECTOR, '.kyber-table .p-1 a') qb_grades = [] qb_names = [] for grade in all_off_grades: qb_grades.append(grade.text) for qb_name in all_qb_names: qb_names.append(qb_name.text) print(qb_grades) print(qb_names)
目标元素结构
评分元素(需提取91.5):
<div class="kyber-grade-badge__info-text">91.5</div>
姓名元素(需提取Josh Allen):
<a class="p-1" href="/nfl/players/2022/REG/josh-allen/46601/passing">Josh Allen</a>
解决方法
修正CSS选择器
- 评分元素的选择器多了一层
div:原选择器.kyber-table .kyber-grade-badge__info-text div会查找该类下的子div,但实际目标文本直接在.kyber-grade-badge__info-text标签内,应改为.kyber-table .kyber-grade-badge__info-text - 姓名选择器可优化为更精准的规则,比如结合父元素定位:
.kyber-table tbody tr td:first-child a.p-1,避免误抓其他无关元素
- 评分元素的选择器多了一层
替换固定等待为显式等待
固定sleep()无法适配页面加载速度,改用WebDriverWait等待元素渲染完成,避免提前查找导致获取不到内容。修改后的代码示例
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC service = Service(executable_path=r"C:\chromedriver.exe") op = webdriver.ChromeOptions() driver = webdriver.Chrome(service=service, options=op) wait = WebDriverWait(driver, 10) # 设置10秒超时 driver.get("https://premium.pff.com/nfl/positions/2022/REG/passing?position=QB") # 等待登录按钮可点击并点击 sign_in = wait.until(EC.element_to_be_clickable((By.XPATH, '/html/body/div/div/header/div[3]/button'))) sign_in.click() # 等待邮箱输入框可见并输入 email = wait.until(EC.visibility_of_element_located((By.XPATH, '/html/body/div/div/div/div/div/div/form/div[1]/input'))) email.send_keys(my_email) # 等待密码输入框可见并输入 password = wait.until(EC.visibility_of_element_located((By.XPATH, '/html/body/div/div/div/div/div/div/form/div[2]/input'))) password.send_keys(my_password) # 等待登录按钮可点击并点击 sign_in_2 = wait.until(EC.element_to_be_clickable((By.XPATH, '/html/body/div/div/div/div/div/div/form/button'))) sign_in_2.click() # 等待表格主体加载完成,再查找目标元素 wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, '.kyber-table tbody'))) # 使用修正后的选择器 all_off_grades = driver.find_elements(By.CSS_SELECTOR, '.kyber-table .kyber-grade-badge__info-text') all_qb_names = driver.find_elements(By.CSS_SELECTOR, '.kyber-table tbody tr td:first-child a.p-1') # 过滤空文本,避免无效数据 qb_grades = [grade.text for grade in all_off_grades if grade.text] qb_names = [name.text for name in all_qb_names if name.text] print(qb_grades) print(qb_names) driver.quit()
额外检查点
- 确认登录后页面是否有弹窗需要关闭,若有需添加对应处理逻辑
- 检查页面是否存在动态分页,若数据分多页,需实现翻页逻辑
内容的提问来源于stack exchange,提问作者Jbucks3
相关产品推荐
相关产品推荐

