Python检索HTML网页表格关键词遇问题,求正确实现方案
问题
需要用Python检索两个网站(https://sslmate.com/labs/crl_watch/ 和 https://sslmate.com/labs/ocsp_watch/)的HTML表格中是否存在特定关键词(如GoDaddy)。尝试两种方法均失败:
- 使用
requests获取页面文本后,通过re.findall和字符串find查找,仅能获取页面头部内容,无法定位表格内的数据; - 用BeautifulSoup解析id为“operator problems”的表格,代码输出对象而非实际表格数据值。
尝试的requests代码
page = requests.get("https://sslmate.com/labs/ocsp_watch/").text print(page) print(re.findall("GoDaddy", page)) print(page.find("GoDaddy"))
尝试的BeautifulSoup代码
from bs4 import BeautifulSoup as bs temp = urllib.request.urlopen('https://sslmate.com/labs/crl_watch/') HTML = temp.read().decode("utf-8") soup = bs(HTML) table = soup.find(table = soup.find("table", attrs={"id":"operator problems"})) headings = [th.get_text() for th in table.find("tr").find_all("th")] datasets = [] for row in table.find_all("tr")[1:]: dataset = headings, (td.get_text() for td in row.find_all("td")) datasets.append(dataset) print(datasets)
原因分析
- 直接用
requests获取文本查找失败:这两个网站的表格数据是动态加载的,初始HTTP请求返回的HTML中不包含表格内容,需等待JavaScript执行渲染后才能获取; - BeautifulSoup代码错误:
soup.find(table = soup.find(...))是无效语法,且静态解析无法获取动态渲染的内容。
正确实现方法
由于表格是动态加载的,需使用支持JavaScript渲染的工具,以下是两种可行方案:
方案1:使用requests-html(无需浏览器驱动)
requests-html可自动执行JavaScript,获取渲染后的页面:
from requests_html import HTMLSession def check_keyword_in_table(url, keyword): session = HTMLSession() r = session.get(url) # 等待页面JavaScript渲染完成 r.html.render() # 定位目标表格 table = r.html.find('#operator problems', first=True) if not table: print(f"未找到目标表格:{url}") return False # 检查关键词是否存在于表格文本中 if keyword in table.text: print(f"在{url}的表格中找到关键词:{keyword}") return True else: print(f"在{url}的表格中未找到关键词:{keyword}") return False # 检测两个目标网站 target_urls = [ "https://sslmate.com/labs/crl_watch/", "https://sslmate.com/labs/ocsp_watch/" ] for url in target_urls: check_keyword_in_table(url, "GoDaddy")
方案2:使用Selenium(需浏览器驱动)
若已安装Selenium及对应浏览器驱动(如ChromeDriver),可使用此方案:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC def check_keyword_with_selenium(url, keyword): # 初始化Chrome浏览器(需确保ChromeDriver已配置到环境变量) driver = webdriver.Chrome() driver.get(url) try: # 等待表格加载,最长等待10秒 table = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "operator problems")) ) # 检查关键词 if keyword in table.text: print(f"在{url}的表格中找到关键词:{keyword}") result = True else: print(f"在{url}的表格中未找到关键词:{keyword}") result = False except Exception as e: print(f"处理{url}时出错:{str(e)}") result = False finally: driver.quit() return result # 检测两个目标网站 target_urls = [ "https://sslmate.com/labs/crl_watch/", "https://sslmate.com/labs/ocsp_watch/" ] for url in target_urls: check_keyword_with_selenium(url, "GoDaddy")
进阶方案:直接调用API接口(效率更高)
打开浏览器开发者工具,查看XHR请求,可找到表格数据的API接口,直接用requests请求接口获取JSON数据,无需渲染页面。示例代码如下:
import requests def check_keyword_via_api(api_url, keyword): response = requests.get(api_url) data = response.json() # 根据接口返回的数据结构,遍历检查关键词 for item in data: if keyword in str(item): print(f"API返回结果中找到关键词:{keyword}") return True print(f"API返回结果中未找到关键词:{keyword}") return False # 替换为实际抓包得到的API地址 # check_keyword_via_api("实际API接口URL", "GoDaddy")
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

