无需Selenium,如何快速实现VISL句法树网页抓取?
优化批量句法树提取的方案
一、直接用HTTP请求替代Selenium(最快方案)
Selenium模拟浏览器渲染页面的开销极大,直接通过requests发送POST请求调用网站接口,跳过浏览器环节,速度能提升一个数量级。
代码示例
import requests from bs4 import BeautifulSoup def get_parse_tree(sentence): target_url = "https://edu.visl.dk/visl/en/parsing/automatic/trees.php" # 匹配页面表单的提交参数 form_data = { "visual": "Vertical", "text": sentence, "lang": "en", "parse": "Parse" } response = requests.post(target_url, data=form_data) soup = BeautifulSoup(response.text, "html.parser") result_block = soup.find("pre") return result_block.get_text(strip=False) if result_block else None # 批量处理示例 sentences = [ "John killed the cat with a hammer.", "She walked to the store quickly.", "The dog chased the ball through the park." ] for sent in sentences: tree = get_parse_tree(sent) print(tree)
二、优化Selenium复用浏览器实例(次优方案)
如果必须依赖Selenium(比如网站有反爬限制),不要每次循环都重启浏览器,复用单个实例能大幅减少启动/关闭的开销。
代码示例
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import Select, WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup # 仅初始化一次浏览器 options = webdriver.ChromeOptions() options.add_argument('--headless') options.add_argument('--disable-gpu') options.add_argument('--no-sandbox') options.add_argument('--blink-settings=imagesEnabled=false') # 禁用图片加载 service = webdriver.ChromeService(executable_path="c:\Program Files (x86)\chromedriver.exe") driver = webdriver.Chrome(service=service, options=options) wait = WebDriverWait(driver, 5) # 仅加载一次页面并设置选项 driver.get("https://edu.visl.dk/visl/en/parsing/automatic/trees.php") form = wait.until(EC.presence_of_element_located((By.NAME, "theform"))) dropdown = form.find_element(By.NAME, "visual") Select(dropdown).select_by_visible_text("Vertical") input_box = form.find_element(By.NAME, "text") submit_btn = form.find_element(By.CSS_SELECTOR, "input[type='submit']") # 循环处理句子 sentences = [ "John killed the cat with a hammer.", "She walked to the store quickly.", "The dog chased the ball through the park." ] for sent in sentences: input_box.clear() input_box.send_keys(sent) submit_btn.click() # 等待结果加载完成 result_block = wait.until(EC.presence_of_element_located((By.TAG_NAME, "pre"))) soup = BeautifulSoup(result_block.get_attribute("innerHTML"), "html.parser") print(soup.get_text(strip=False)) # 返回表单页面,避免重复加载整页 driver.back() form = wait.until(EC.presence_of_element_located((By.NAME, "theform"))) input_box = form.find_element(By.NAME, "text") submit_btn = form.find_element(By.CSS_SELECTOR, "input[type='submit']") driver.quit()
三、额外优化点
- 用显式等待
WebDriverWait替代硬等待,确保元素加载完成再操作,减少无效等待时间 - 如果网站支持批量输入,可将多个句子用换行分隔后一次性提交,减少请求次数
内容的提问来源于stack exchange,提问作者Mufarrid Ansari
相关产品推荐
相关产品推荐

