You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Perl脚本读取CSV文件中货币符号乱码问题求助

解决Perl读写CSV时货币符号乱码问题

从你提供的十六进制数据能明确:输入文件是UTF-8编码:

  • c2 a3 是UTF-8格式的英镑符号£(对应Unicode U+00A3)
  • e2 82 ac 是UTF-8格式的欧元符号€(对应Unicode U+20AC)

你之前用>:encoding(latin1)输出的问题在于:Latin1(ISO-8859-1)编码里根本没有欧元符号——€是Windows-1252编码特有的字符,对应字节0x80。Perl遇到无法用Latin1编码的字符(比如U+20AC)时,会自动把它转义成\x{20ac}的字符串形式,这就是你输出里出现\x{20ac}的原因。

修正方案

把输出编码改成Windows-1252(cp1252),这是你期望的编码(因为目标€对应0x80)。修正后的代码如下:

open my $in,  '<:encoding(UTF-8)',  'input-file-name'  or die $!;
open my $out, '>:encoding(cp1252)', 'output-file-name' or die $!;

while ( <$in> ) {
   print $out $_;
}

# 关闭文件句柄(可选,但属于良好编程习惯)
close $in;
close $out;

效果说明

  • 读取阶段:<:encoding(UTF-8)把输入的UTF-8字节解码成Perl内部的Unicode字符串,确保字符在内存中是正确的。
  • 输出阶段::encoding(cp1252)把Unicode字符串编码成Windows-1252字节流:
    • £(U+00A3)对应cp1252的0xA3,和你期望的输出一致。
    • €(U+20AC)对应cp1252的0x80,完美输出目标符号。

修正后输出的十六进制应该为:

$ od -t x1 output-file-name
0000000 47 65 74 20 a3 35 30 20 6f 72 20 80 35 30 20 64
0000020 61 69 6c 79 0a

内容的提问来源于stack exchange,提问作者Developer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 06:40:28