You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取求助:提取样式表与WOFF文件中的希腊/希伯来文本

解决方案:抓取浏览器渲染后的特殊文本(希腊文/希伯来文)

这两个页面的问题核心是文本通过自定义字体映射渲染:DOM中存储的是占位符或普通字符,浏览器通过自定义字体文件将其替换为希腊/希伯来文,但基础爬虫工具只能拿到原始DOM内容,无法获取渲染后的实际文本。以下是几种可行的解决方法:

方法1:利用Chrome DevTools Protocol(CDP)直接获取渲染后文本

基础Selenium获取的是DOM原始内容,而CDP可以直接拿到浏览器渲染完成后的页面文本,绕过字体映射的问题。

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# 初始化Chrome浏览器(无头模式可选)
options = Options()
options.add_argument("--headless=new")  # 无头模式,无需显示浏览器窗口
driver = webdriver.Chrome(options=options)

# 抓取希腊文本页面
driver.get("https://777codes.com/newtestament/mat1.html")
greek_text = driver.execute_cdp_cmd(
    "Runtime.evaluate",
    {"expression": "document.body.innerText"}
)["result"]["value"]
print("希腊文本内容:\n", greek_text)

# 抓取希伯来文本页面
driver.get("https://777codes.com/newtestament/gen1.html")
hebrew_text = driver.execute_cdp_cmd(
    "Runtime.evaluate",
    {"expression": "document.body.innerText"}
)["result"]["value"]
print("\n希伯来文本内容:\n", hebrew_text)

driver.quit()

方法2:通过剪贴板复制渲染后文本

利用浏览器的复制功能,直接复制页面渲染完成的文本,再读取剪贴板内容,这种方法对字体映射类的文本抓取非常有效。

from selenium import webdriver
import pyperclip

driver = webdriver.Chrome()

# 处理希伯来文本页面
driver.get("https://777codes.com/newtestament/gen1.html")
# 选中页面全部内容并复制
driver.execute_script("document.body.focus(); document.execCommand('SelectAll'); document.execCommand('Copy');")
# 读取剪贴板内容
hebrew_text = pyperclip.paste()
print("希伯来文本:\n", hebrew_text)

# 处理希腊文本页面同理
driver.get("https://777codes.com/newtestament/mat1.html")
driver.execute_script("document.body.focus(); document.execCommand('SelectAll'); document.execCommand('Copy');")
greek_text = pyperclip.paste()
print("\n希腊文本:\n", greek_text)

driver.quit()

方法3:解析自定义字体映射(进阶)

如果上面两种方法失效,可以直接解析网站的自定义字体文件,建立占位字符与目标字符的映射关系:

  1. 打开页面,在Chrome开发者工具的Sources面板找到字体文件(从CSS的@font-face规则中获取URL)
  2. 下载字体文件(通常是.woff2或.ttf格式)
  3. 使用Python的fonttools库解析字体,提取glyph与Unicode字符的对应关系
  4. 将DOM中获取的占位字符替换为对应的希腊/希伯来字符

示例代码(需先安装fonttools:pip install fonttools):

from fontTools.ttLib import TTFont

# 下载字体文件后解析
font = TTFont("custom_font.woff2")
# 获取字符到glyph的映射
cmap = font.getBestCmap()
# 反转映射,得到glyph到字符的对应关系
glyph_to_char = {v: k for k, v in cmap.items()}
# 之后将DOM中获取的占位字符通过映射转换为目标文本

内容的提问来源于stack exchange,提问作者ShaneO

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 02:00:38