You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需Selenium,如何快速实现VISL句法树网页抓取?

优化批量句法树提取的方案

一、直接用HTTP请求替代Selenium(最快方案)

Selenium模拟浏览器渲染页面的开销极大,直接通过requests发送POST请求调用网站接口,跳过浏览器环节,速度能提升一个数量级。

代码示例

import requests
from bs4 import BeautifulSoup

def get_parse_tree(sentence):
    target_url = "https://edu.visl.dk/visl/en/parsing/automatic/trees.php"
    # 匹配页面表单的提交参数
    form_data = {
        "visual": "Vertical",
        "text": sentence,
        "lang": "en",
        "parse": "Parse"
    }
    response = requests.post(target_url, data=form_data)
    soup = BeautifulSoup(response.text, "html.parser")
    result_block = soup.find("pre")
    return result_block.get_text(strip=False) if result_block else None

# 批量处理示例
sentences = [
    "John killed the cat with a hammer.",
    "She walked to the store quickly.",
    "The dog chased the ball through the park."
]

for sent in sentences:
    tree = get_parse_tree(sent)
    print(tree)

二、优化Selenium复用浏览器实例(次优方案)

如果必须依赖Selenium(比如网站有反爬限制),不要每次循环都重启浏览器,复用单个实例能大幅减少启动/关闭的开销。

代码示例

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 仅初始化一次浏览器
options = webdriver.ChromeOptions()
options.add_argument('--headless')
options.add_argument('--disable-gpu')
options.add_argument('--no-sandbox')
options.add_argument('--blink-settings=imagesEnabled=false')  # 禁用图片加载
service = webdriver.ChromeService(executable_path="c:\Program Files (x86)\chromedriver.exe")
driver = webdriver.Chrome(service=service, options=options)
wait = WebDriverWait(driver, 5)

# 仅加载一次页面并设置选项
driver.get("https://edu.visl.dk/visl/en/parsing/automatic/trees.php")
form = wait.until(EC.presence_of_element_located((By.NAME, "theform")))
dropdown = form.find_element(By.NAME, "visual")
Select(dropdown).select_by_visible_text("Vertical")
input_box = form.find_element(By.NAME, "text")
submit_btn = form.find_element(By.CSS_SELECTOR, "input[type='submit']")

# 循环处理句子
sentences = [
    "John killed the cat with a hammer.",
    "She walked to the store quickly.",
    "The dog chased the ball through the park."
]

for sent in sentences:
    input_box.clear()
    input_box.send_keys(sent)
    submit_btn.click()
    # 等待结果加载完成
    result_block = wait.until(EC.presence_of_element_located((By.TAG_NAME, "pre")))
    soup = BeautifulSoup(result_block.get_attribute("innerHTML"), "html.parser")
    print(soup.get_text(strip=False))
    # 返回表单页面,避免重复加载整页
    driver.back()
    form = wait.until(EC.presence_of_element_located((By.NAME, "theform")))
    input_box = form.find_element(By.NAME, "text")
    submit_btn = form.find_element(By.CSS_SELECTOR, "input[type='submit']")

driver.quit()

三、额外优化点

  • 用显式等待WebDriverWait替代硬等待,确保元素加载完成再操作,减少无效等待时间
  • 如果网站支持批量输入,可将多个句子用换行分隔后一次性提交,减少请求次数

内容的提问来源于stack exchange,提问作者Mufarrid Ansari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 20:57:05