You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python检索HTML网页表格关键词遇问题,求正确实现方案

问题

需要用Python检索两个网站(https://sslmate.com/labs/crl_watch/ 和 https://sslmate.com/labs/ocsp_watch/)的HTML表格中是否存在特定关键词(如GoDaddy)。尝试两种方法均失败:

  1. 使用requests获取页面文本后,通过re.findall和字符串find查找,仅能获取页面头部内容,无法定位表格内的数据;
  2. 用BeautifulSoup解析id为“operator problems”的表格,代码输出对象而非实际表格数据值。

尝试的requests代码

page = requests.get("https://sslmate.com/labs/ocsp_watch/").text
print(page)
print(re.findall("GoDaddy", page))
print(page.find("GoDaddy"))

尝试的BeautifulSoup代码

from bs4 import BeautifulSoup as bs 

temp = urllib.request.urlopen('https://sslmate.com/labs/crl_watch/')
HTML = temp.read().decode("utf-8")
soup = bs(HTML)
table = soup.find(table = soup.find("table", attrs={"id":"operator problems"}))
headings = [th.get_text() for th in table.find("tr").find_all("th")]

datasets = []
for row in table.find_all("tr")[1:]:
    dataset = headings, (td.get_text() for td in row.find_all("td"))
    datasets.append(dataset)
  
print(datasets)

原因分析

  • 直接用requests获取文本查找失败:这两个网站的表格数据是动态加载的,初始HTTP请求返回的HTML中不包含表格内容,需等待JavaScript执行渲染后才能获取;
  • BeautifulSoup代码错误:soup.find(table = soup.find(...))是无效语法,且静态解析无法获取动态渲染的内容。

正确实现方法

由于表格是动态加载的,需使用支持JavaScript渲染的工具,以下是两种可行方案:

方案1:使用requests-html(无需浏览器驱动)

requests-html可自动执行JavaScript,获取渲染后的页面:

from requests_html import HTMLSession

def check_keyword_in_table(url, keyword):
    session = HTMLSession()
    r = session.get(url)
    # 等待页面JavaScript渲染完成
    r.html.render()
    
    # 定位目标表格
    table = r.html.find('#operator problems', first=True)
    if not table:
        print(f"未找到目标表格:{url}")
        return False
    
    # 检查关键词是否存在于表格文本中
    if keyword in table.text:
        print(f"在{url}的表格中找到关键词:{keyword}")
        return True
    else:
        print(f"在{url}的表格中未找到关键词:{keyword}")
        return False

# 检测两个目标网站
target_urls = [
    "https://sslmate.com/labs/crl_watch/",
    "https://sslmate.com/labs/ocsp_watch/"
]
for url in target_urls:
    check_keyword_in_table(url, "GoDaddy")

方案2:使用Selenium(需浏览器驱动)

若已安装Selenium及对应浏览器驱动(如ChromeDriver),可使用此方案:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def check_keyword_with_selenium(url, keyword):
    # 初始化Chrome浏览器(需确保ChromeDriver已配置到环境变量)
    driver = webdriver.Chrome()
    driver.get(url)
    
    try:
        # 等待表格加载,最长等待10秒
        table = WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.ID, "operator problems"))
        )
        # 检查关键词
        if keyword in table.text:
            print(f"在{url}的表格中找到关键词:{keyword}")
            result = True
        else:
            print(f"在{url}的表格中未找到关键词:{keyword}")
            result = False
    except Exception as e:
        print(f"处理{url}时出错:{str(e)}")
        result = False
    finally:
        driver.quit()
    return result

# 检测两个目标网站
target_urls = [
    "https://sslmate.com/labs/crl_watch/",
    "https://sslmate.com/labs/ocsp_watch/"
]
for url in target_urls:
    check_keyword_with_selenium(url, "GoDaddy")

进阶方案:直接调用API接口(效率更高)

打开浏览器开发者工具,查看XHR请求,可找到表格数据的API接口,直接用requests请求接口获取JSON数据,无需渲染页面。示例代码如下:

import requests

def check_keyword_via_api(api_url, keyword):
    response = requests.get(api_url)
    data = response.json()
    # 根据接口返回的数据结构,遍历检查关键词
    for item in data:
        if keyword in str(item):
            print(f"API返回结果中找到关键词:{keyword}")
            return True
    print(f"API返回结果中未找到关键词:{keyword}")
    return False

# 替换为实际抓包得到的API地址
# check_keyword_via_api("实际API接口URL", "GoDaddy")

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 06:30:39