使用Selenium和BeautifulSoup爬取网页表格时无法获取<tr>元素的问题
爬取Leetify比赛详情页时无法获取
<tbody>元素的原因与解决办法 问题描述
我尝试用Selenium+BeautifulSoup爬取Leetify比赛详情页(https://leetify.com/app/match-details/5c438e85-c31c-443a-8257-5872d89e548c/details-general)的表格,但调用BeautifulSoup.find_all("tbody")返回空数组。浏览器审查元素能看到多个<tbody>和<tr>元素,但代码就是获取不到。我的代码如下:
from selenium import webdriver from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get("https://leetify.com/app/match-details/5c438e85-c31c-443a-8257-5872d89e548c/details-general") html_source = driver.page_source soup = BeautifulSoup(html_source, 'html.parser') table = soup.find_all("tbody") print(len(table)) for entry in table: print(entry) print("\n")
原因分析
- 页面动态加载未完成:
driver.get()仅等待页面初始HTML加载完成,但表格这类数据往往是通过AJAX异步请求获取后渲染的,此时driver.page_source拿到的源码还没有包含表格的DOM元素。 - 未登录导致内容未加载:Leetify的比赛详情页需要登录才能查看完整数据,未登录状态下页面只会渲染框架,不会加载实际的表格内容。
- iframe嵌套(可能性较低):少数网站会把内容放在iframe中,直接获取页面源码无法拿到iframe内部的DOM,但该页面大概率不存在这个问题。
解决方法
1. 等待动态元素加载完成
使用Selenium的显式等待,指定等待目标<tbody>元素出现后再获取页面源码,确保内容已渲染:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException from bs4 import BeautifulSoup driver = webdriver.Chrome() driver.get("https://leetify.com/app/match-details/5c438e85-c31c-443a-8257-5872d89e548c/details-general") try: # 最多等待10秒,直到至少一个tbody元素出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "tbody")) ) except TimeoutException: print("等待表格加载超时") driver.quit() html_source = driver.page_source soup = BeautifulSoup(html_source, 'html.parser') table = soup.find_all("tbody") print(len(table)) for entry in table: print(entry) print("\n") driver.quit()
2. 处理登录验证
如果页面需要登录才能查看内容,可添加登录逻辑,或手动登录后继续爬取:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup driver = webdriver.Chrome() # 先跳转到登录页,等待用户手动登录 driver.get("https://leetify.com/login") input("请在浏览器中完成登录,登录后按回车键继续...") # 登录完成后再访问目标页面 driver.get("https://leetify.com/app/match-details/5c438e85-c31c-443a-8257-5872d89e548c/details-general") # 等待表格加载 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.TAG_NAME, "tbody")) ) html_source = driver.page_source soup = BeautifulSoup(html_source, 'html.parser') table = soup.find_all("tbody") print(len(table)) for entry in table: print(entry) print("\n") driver.quit()
3. 直接用Selenium提取元素
如果BeautifulSoup仍无法解析,可直接用Selenium定位并提取表格内容,无需转换为BeautifulSoup对象:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("https://leetify.com/app/match-details/5c438e85-c31c-443a-8257-5872d89e548c/details-general") # 等待所有tbody元素加载完成 tbodies = WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.TAG_NAME, "tbody")) ) # 遍历输出每个tbody的内容 for idx, tbody in enumerate(tbodies): print(f"第{idx+1}个tbody内容:") print(tbody.get_attribute("innerHTML")) print("\n") driver.quit()
内容的提问来源于stack exchange,提问作者Horde Bob
相关产品推荐
相关产品推荐

