Mule 4 DataWeave处理CSV:UTF-8标点转换与换行清理问题
Mule 4 DataWeave 处理CSV数据问题解决方案
问题说明
需要处理包含内嵌换行的CSV数据(行终止符为换行符),核心需求有两个:
- 将UTF-8标点转换为对应ASCII标点
- 移除ASCII 00-7F范围之外的所有字符
但当前的DataWeave代码里,标点转换逻辑未生效,内嵌换行的清理也没起作用。
现有代码
%dw 2.0 output application/csv fun cleanText(text: String): String = text replace /\r?\n/ with " " // 尝试清理内嵌换行 replace /[“”‘’]/ with "'" // 尝试转换UTF-8引号 replace /[–—]/ with "-" // 尝试转换UTF-8破折号 replace /[…]/ with "..." // 尝试转换UTF-8省略号 replace /[^\x00-\x7F]/ with "" // 尝试移除非ASCII字符 --- payload map ($ mapObject (value, key) -> { (key): cleanText(value as String) } )
样例数据
id,description,notes 1,"This is a test with “quoted text” and a line break inside",Some notes with —em dash… and non-ASCII char: é 2,Another entry without issues,Plain text here
预期输出
id,description,notes 1,"This is a test with 'quoted text' and a line break inside",Some notes with -em dash... and non-ASCII char: 2,Another entry without issues,Plain text here
实际输出
id,description,notes 1,"This is a test with “quoted text” and a line break inside",Some notes with —em dash… and non-ASCII char: é 2,Another entry without issues,Plain text here
修复方案
1. 问题根源
- 内嵌换行未清理:原正则仅匹配
\r?\n,未覆盖所有换行类型,且CSV解析后字段内的换行符需明确匹配替换 - 标点转换失效:直接写入UTF-8标点字符可能因编码匹配问题无法命中,需改用Unicode转义序列
修改后的完整代码
%dw 2.0 output application/csv quoteValues=true // 确保输出字段格式符合CSV规范 fun cleanText(text: String): String = text // 替换所有类型的换行符为空格 replace /\r\n|\r|\n/ with " " // 用Unicode转义匹配UTF-8标点,避免编码问题 replace /\u201C|\u201D|\u2018|\u2019/ with "'" // 匹配左右双/单引号 replace /\u2013|\u2014/ with "-" // 匹配短/长破折号 replace /\u2026/ with "..." // 匹配省略号 // 移除所有非ASCII字符 replace /[^\x00-\x7F]/ with "" --- payload map ($ mapObject (value, key) -> { (key): if (value is String) cleanText(value) else value } )
关键修改点
- 扩展换行符匹配规则,覆盖
\r\n、\r、\n所有场景 - 使用Unicode转义序列替代直接写入UTF-8字符,确保匹配准确
- 输出时启用
quoteValues=true,保证处理后的带空格字段正确用引号包裹 - 增加类型判断,避免非字符串类型字段触发报错
内容的提问来源于stack exchange,提问作者ramesh.metta
相关产品推荐
相关产品推荐

