You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将SarcasmAmazonReviewsCorpus数据集的HTML内容转为JSON?

将HTML格式的亚马逊讽刺评论数据集转换为JSON

可以通过解析HTML提取目标字段再序列化到JSON文件的方式完成,具体步骤和代码如下:

1. 导入依赖库

需要用到文件操作、HTML解析和JSON处理的相关库:

import os
from bs4 import BeautifulSoup
import json

2. 读取所有TXT文件内容

优化读取逻辑,适配旧数据集可能存在的编码问题:

path = "你的数据集存放路径"
files_content = []

for filename in os.listdir(path):
    if filename.endswith(".txt"):
        filepath = os.path.join(path, filename)
        # 旧数据集常用latin-1编码,避免读取报错
        with open(filepath, mode='r', encoding='latin-1') as f:
            files_content.append(f.read())

3. 解析HTML并提取字段

根据数据集的HTML结构,提取评论、产品名称、互动信息等字段(需根据实际标签调整选择器):

parsed_reviews = []

for content in files_content:
    soup = BeautifulSoup(content, 'html5lib')
    review_item = {}
    
    # 提取产品名称(示例:假设产品名在<h2>标签内)
    product_tag = soup.find('h2')
    if product_tag:
        review_item["product_name"] = product_tag.get_text(strip=True)
    
    # 提取评论内容(示例:假设评论在class为review-text的div内)
    review_text_tag = soup.find('div', class_='review-text')
    if review_text_tag:
        review_item["review_text"] = review_text_tag.get_text(strip=True)
    
    # 提取互动信息(示例:假设有用数在class为helpful的span内)
    helpful_tag = soup.find('span', class_='helpful')
    if helpful_tag:
        review_item["helpful_count"] = helpful_tag.get_text(strip=True)
    
    # 可根据实际结构添加更多字段,比如评分、评论日期等
    parsed_reviews.append(review_item)

4. 写入JSON文件

将解析后的结构化数据写入JSON文件:

with open('amazon_sarcasm_reviews.json', 'w', encoding='utf-8') as f:
    json.dump(parsed_reviews, f, indent=4, ensure_ascii=False)

注意事项

  • 必须根据你手中数据集的实际HTML结构调整标签选择器(比如class名称、标签类型),不同版本的数据集结构可能有差异
  • 如果读取时出现编码错误,可尝试替换encoding参数为utf-8、cp1252等
  • 若部分字段提取失败,可添加默认值处理,比如review_item["product_name"] = product_tag.get_text(strip=True) if product_tag else "未知"

内容的提问来源于stack exchange,提问作者Guilherme Fulop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 07:05:19