You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python删除JSON文件中name字段含泰语等非英文内容的行

实现方案

核心逻辑是通过泰语对应的Unicode编码区间判断字符是否为泰语,逐行读取JSONL格式的文件,过滤name字段含泰语的行后写入新文件。

完整可运行代码

import json

def has_thai_char(s: str) -> bool:
    """判断字符串是否包含泰语字符"""
    # 泰语Unicode编码范围 U+0E00 ~ U+0E7F
    for char in s:
        if '\u0e00' <= char <= '\u0e7f':
            return True
    return False

if __name__ == '__main__':
    # 按需修改输入输出文件路径
    input_file = 'data.json'
    output_file = 'filtered_data.json'

    with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f:
        for line in in_f:
            stripped_line = line.strip()
            # 跳过空行
            if not stripped_line:
                continue
            # 解析JSON
            try:
                data = json.loads(stripped_line)
            except json.JSONDecodeError:
                # 格式非法的行默认跳过,如有需要可以改为写入输出文件
                continue
            # 过滤掉name含泰语的行
            if not has_thai_char(data.get('name', '')):
                out_f.write(line)

注意事项

  • 你提供的输入示例中car字段的值(如audi、mercedes)没有加双引号,属于非法JSON格式,实际使用前请保证每行JSON格式合法,否则会被异常捕获逻辑跳过。
  • 建议处理完后校验输出文件符合预期,再决定是否覆盖原文件。
  • 代码默认使用utf-8编码读写文件,避免出现泰语乱码问题。

内容的提问来源于stack exchange,提问作者sirimiri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 18:15:02