You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas to_csv保存CSV时GBK编码报错的技术咨询

解决CSV保存时的GBK编码错误

错误原因

\u2022 是Unicode中的圆点符号(•),GBK编码集不包含这个字符,所以用encoding='gbk'保存数据时会触发编码失败。

解决方案

方案1:改用UTF-8编码(推荐)

UTF-8支持所有Unicode字符,为了让Excel能正常识别内容,建议使用带BOM的UTF-8编码:

data.to_csv('/home/bio_kang/Learning/Python/film_project/top250_film_info.csv', index=None, encoding='utf-8-sig')

方案2:保留GBK编码,处理异常字符

如果必须用GBK编码,可以通过errors参数忽略或替换无法编码的字符:

# 忽略无法编码的字符
data.to_csv('/home/bio_kang/Learning/Python/film_project/top250_film_info.csv', index=None, encoding='gbk', errors='ignore')

# 用问号替换无法编码的字符
data.to_csv('/home/bio_kang/Learning/Python/film_project/top250_film_info.csv', index=None, encoding='gbk', errors='replace')

方案3:提前清理特殊字符

手动替换GBK不支持的字符,比如把\u2022换成GBK兼容的中文圆点·:

# 以简介列为例替换特殊字符,其他列可同理处理
data['description'] = data['description'].str.replace('\u2022', '·')
# 处理后再保存
data.to_csv('/home/bio_kang/Learning/Python/film_project/top250_film_info.csv', index=None, encoding='gbk')

额外代码优化建议

你的正则表达式存在一些小问题,可能导致数据提取不准确:

  • 导演提取正则:原写法缺少content后的引号,修正后:
    pattern_director = re.compile(r'<meta content="(.*?)" property="video:director"/>')
    
  • 简介提取正则:原写法缺少content值后的引号,修正后:
    pattern_description = re.compile(r'<meta content="(.*?)" property="og:description">')
    
  • 提取列表数据时,不要用str(re.findall(...))转成带方括号的字符串,建议直接取匹配结果的第一个元素(确保有结果时):
    # 年份提取示例,替换原写法
    year_match = re.findall(pattern_year, film_info)
    years.append(year_match[0] if year_match else '')
    

内容的提问来源于stack exchange,提问作者kang Ouyang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 14:51:23