You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多语言CSV读取与翻译优化技术问题求助

解决CSV多语言读取与翻译优化问题

一、修复多语言编码乱码问题

  • 原代码用latin1编码读取,它是单字节编码,无法兼容中文、韩文等多字节语言,必然出现乱码。推荐使用**utf-8-sig**编码,它能处理带BOM的UTF-8文件,几乎兼容所有主流语言。
  • 修改文件读取部分的编码参数:
    with open(file_name, 'r', encoding='utf-8-sig') as myfile:
    
  • 如果不确定文件编码,可以用chardet库自动检测:
    1. 先安装依赖:pip install chardet
    2. 检测编码的代码片段:
      import chardet
      with open(file_name, 'rb') as f:
          result = chardet.detect(f.read())
      # 用检测出的编码打开文件
      with open(file_name, 'r', encoding=result['encoding']) as myfile:
      

二、跳过英文评论提升翻译性能

  • 借助langdetect库检测评论语言,仅翻译非英文内容:
    1. 安装依赖:pip install langdetect
    2. 在代码中加入语言检测逻辑,跳过英文翻译:
      from langdetect import detect, LangDetectException
      
      处理评论时的逻辑修改:
      if len(comments) > 10:
          try:
              if detect(comments) != 'en':
                  translated = tss.google(comments)
                  group_sentences[listing_id].append(translated)
              else:
                  # 直接保留原英文评论
                  group_sentences[listing_id].append(comments)
          except LangDetectException:
              # 无法检测语言时,默认执行翻译
              group_sentences[listing_id].append(tss.google(comments))
      

完整优化后代码

import pandas as pd
from collections import defaultdict
import csv
import translators.server as tss
from langdetect import detect, LangDetectException
import chardet

group_sentences = {}

# 自动检测文件编码(也可直接替换为utf-8-sig)
with open(file_name, 'rb') as f:
    detect_result = chardet.detect(f.read())

with open(file_name, 'r', encoding=detect_result['encoding']) as myfile: 
    reader = csv.reader(myfile, delimiter=';')
    for n, row in enumerate(reader):
        if not n:
            continue  
        listing_id, id, date, reviewer_id, reviewer_name, comments = row
        if listing_id not in group_sentences:
            group_sentences[listing_id] = list()
        if len(comments) > 10:
            try:
                # 跳过英文评论,直接保留原文
                if detect(comments) == 'en':
                    group_sentences[listing_id].append(comments)
                else:
                    # 翻译非英文评论
                    translated_comment = tss.google(comments)
                    group_sentences[listing_id].append(translated_comment)
            except LangDetectException:
                # 无法识别语言时,执行翻译
                group_sentences[listing_id].append(tss.google(comments))

内容的提问来源于stack exchange,提问作者Fernando

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:05:17