如何用Python+Beautiful Soup抓取脚本生成的Handlebars表格?
当然可以抓取到表格数据!
topics-template标签是关键线索 首先明确:Handlebars模板本身只是结构框架,但它能帮你定位到实际数据的形态,而你要找的表格数据,大概率藏在页面的其他地方——要么是页面内嵌的JSON,要么是通过AJAX请求加载的接口数据。
为什么topics-template有用?
这个script标签里的内容就是Handlebars的模板代码,它定义了表格的渲染规则。比如你可能会看到类似这样的片段:
{{#each topics}} <tr> <td>{{title}}</td> <td>{{author.name}}</td> <td>{{createdAt}}</td> </tr> {{/each}}
这段代码告诉你:
- 表格的数据来自一个叫
topics的数组 - 每个数组元素包含
title、author.name、createdAt这些字段
这会帮你快速识别要抓取的数据结构,不用瞎猜字段名。
具体抓取步骤
1. 提取模板,明确数据结构
先用Beautiful Soup把模板内容抓出来,分析里面的变量:
from bs4 import BeautifulSoup import requests target_url = "你的目标网站URL" response = requests.get(target_url) soup = BeautifulSoup(response.text, "html.parser") # 获取模板内容 template_tag = soup.find("script", id="topics-template") if template_tag: template_content = template_tag.string print(template_content)
2. 寻找实际数据的来源
有两种常见情况:
- 情况一:数据内嵌在页面的其他script标签里
很多网站会把渲染模板需要的JSON数据直接放在页面的script标签中,比如类似var topicsData = [...];或者window.__DATA__ = {...}的形式。你可以用正则表达式提取:
import re import json # 遍历所有script标签,查找包含数据的部分 for script in soup.find_all("script"): if not script.string: continue # 这里的正则要根据你看到的变量名调整,比如上面模板里的topics match = re.search(r"var topics = (\[.*?\]);", script.string, re.DOTALL) if match: raw_data = match.group(1) topics_data = json.loads(raw_data) # 现在你拿到了原始数据,可以直接处理 for topic in topics_data: print(f"标题: {topic['title']}, 作者: {topic['author']['name']}")
- 情况二:数据通过AJAX接口加载
如果页面里找不到内嵌的JSON,就打开浏览器开发者工具(F12),切换到「Network」标签,刷新页面,过滤「XHR/Fetch」类型的请求。观察每个请求的响应内容,找到返回topics数组的那个接口。找到后,直接用requests请求这个接口就能拿到数据:
# 假设找到的接口URL是这个 api_url = "https://目标网站.com/api/topics" api_response = requests.get(api_url) topics_data = api_response.json() # 处理数据 for topic in topics_data: print(f"标题: {topic['title']}, 发布时间: {topic['createdAt']}")
总结
topics-template标签本身没有数据,但它是你理解数据结构的钥匙。通过它你能知道要找什么字段,然后从页面内嵌的script标签或者AJAX接口中提取原始数据,这样就不用依赖前端渲染,直接高效地拿到你需要的表格内容。
内容的提问来源于stack exchange,提问作者gutscdav000
相关产品推荐
相关产品推荐

