You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取head内JSON-LD中的datePublished字段

提取方案

你定位到目标script标签后,不需要写正则做文本匹配,直接用Python标准库json解析标签内的结构化内容即可,步骤如下:

  • 取出目标script标签的内部文本
  • 调用json.loads()将JSON格式的文本转为Python字典/列表
  • 从解析后的结构化数据中读取datePublished字段即可
完整代码示例
import json
from bs4 import BeautifulSoup

# 初始化BeautifulSoup对象、定位目标script标签的逻辑可直接复用你已写的代码
soup = BeautifulSoup(页面源码字符串, "html.parser")
# 定位所有type为application/ld+json的script标签,避免单个标签不含目标字段的问题
ld_json_scripts = soup.find_all("script", type="application/ld+json")

date_published = None
for script in ld_json_scripts:
    try:
        data = json.loads(script.string)
        # 处理单JSON对象的情况
        if isinstance(data, dict) and data.get("datePublished"):
            date_published = data["datePublished"]
            break
        # 处理JSON数组包裹多组结构化数据的情况
        if isinstance(data, list):
            for item in data:
                if isinstance(item, dict) and item.get("datePublished"):
                    date_published = item["datePublished"]
                    break
            if date_published:
                break
    except json.JSONDecodeError:
        # 遇到格式损坏的JSON块直接跳过,继续检查下一个
        continue

print(date_published)
注意事项
  • 不要用正则表达式直接从HTML文本中匹配datePublished值,遇到JSON转义字符、嵌套结构时很容易出现匹配错误,用标准JSON库解析的稳定性最高
  • 部分页面会放置多个application/ld+json类型的script标签,分别存储文章、面包屑、站点信息等不同结构化数据,用find_all遍历排查比只取第一个标签的容错率更高
  • 代码里加了JSON解析异常捕获,遇到格式不规范的JSON块不会直接中断程序

内容的提问来源于stack exchange,提问作者Luís Henrique Martins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 21:01:58