You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python移除XML特殊字符及CSV导出逗号分隔问题求助

解决CSV字段拆分与XML特殊字符清理问题

咱们一步步来解决你的两个问题:先是CSV里逗号导致内容拆分的问题,然后是XML特殊字符的清理。

一、修复CSV中逗号导致的内容拆分问题

你现在遇到的CSV拆分问题,根源是手动拼接字符串写入文件——这种方式完全绕开了csv模块的内置处理逻辑,自然会让带逗号的内容被错误拆分。

csv模块本身就专门处理这种场景:当字段包含逗号、引号或者换行符时,它会自动用双引号包裹该字段,确保内容完整保留。只需要修改你的save_csv函数,使用writerow()方法逐行写入即可:

def save_csv():
    # 添加newline=''是csv模块的推荐写法,避免跨平台换行符混乱
    with open('titles.csv', 'w', newline='', encoding='utf-8') as f:
        f_csv = csv.writer(f)
        f_csv.writerow(headers)
        # 遍历配对后的列表,用writerow写入每一行
        for t, c, u, p in zip(title, category, url, prod):
            f_csv.writerow([t, c, u, p])

这样修改后,像Add, Change, or Remove Transitions between Slides这类带逗号的标题,会被自动包裹在双引号里写入CSV,打开时就不会被拆分成多个单元格了。

如果你确实有需求要移除标题中的逗号(不推荐,会破坏原始内容),可以在提取标题时用replace方法:

title.append(t.find('title').text.replace(',', ''))

但显然,用csv模块的原生方法保留原内容是更合理的方案。

二、清理XML中的特殊字符

XML里的特殊字符主要分两种情况:

  1. XML实体引用:比如&amp;(对应&)、&lt;(对应<)、&gt;(对应>)等,xml.etree.ElementTree在解析XML时会自动把这些实体转换成对应的正常字符,不需要额外处理。
  2. 未转义的无效字符:比如XML中直接出现<(而不是&lt;),或者一些非打印控制字符,这类情况会导致解析报错,或者提取出的文本有残留垃圾字符。

我们可以写一个清理函数来处理这些情况:

import re

def clean_special_chars(text):
    if text is None:
        return ''
    # 替换XML实体引用为对应正常字符
    text = re.sub(r'&amp;', '&', text)
    text = re.sub(r'&lt;', '<', text)
    text = re.sub(r'&gt;', '>', text)
    text = re.sub(r'&quot;', '"', text)
    text = re.sub(r'&apos;', "'", text)
    # 移除XML不允许的控制字符(可选,根据你的需求调整)
    text = re.sub(r'[\x00-\x08\x0B\x0C\x0E-\x1F]', '', text)
    return text

然后在提取文本的时候调用这个函数,比如:

title.append(clean_special_chars(t.find('title').text))

整合后的完整代码

我还优化了你的数据提取逻辑,减少了重复遍历XML节点的次数,同时添加了空值判断,避免某个标签不存在时抛出错误:

import xml.etree.ElementTree as ET
import csv
import re

def clean_special_chars(text):
    if text is None:
        return ''
    # 解析XML实体引用
    text = re.sub(r'&amp;', '&', text)
    text = re.sub(r'&lt;', '<', text)
    text = re.sub(r'&gt;', '>', text)
    text = re.sub(r'&quot;', '"', text)
    text = re.sub(r'&apos;', "'", text)
    # 移除无效控制字符
    text = re.sub(r'[\x00-\x08\x0B\x0C\x0E-\x1F]', '', text)
    return text

tree = ET.parse('file.xml')
root = tree.getroot()

title = []
category = []
url = []
prod = []

def extract_data():
    # 一次遍历所有solution节点,避免重复遍历
    for solution in root.findall('solution'):
        # 处理head中的title
        head = solution.find('head')
        title_elem = head.find('title') if head is not None else None
        title.append(clean_special_chars(title_elem.text if title_elem is not None else ''))
        
        # 处理body中的各个字段
        body = solution.find('body')
        if body is not None:
            cat_elem = body.find('category')
            category.append(clean_special_chars(cat_elem.text if cat_elem is not None else ''))
            
            video_elem = body.find('video')
            url.append(clean_special_chars(video_elem.text if video_elem is not None else ''))
            
            prod_elem = body.find('product')
            prod.append(clean_special_chars(prod_elem.text if prod_elem is not None else ''))
        else:
            # 如果body不存在,添加空字符串占位
            category.append('')
            url.append('')
            prod.append('')

extract_data()

headers = ['Title', 'Category', 'Video URL','Product']

def save_csv():
    with open('titles.csv', 'w', newline='', encoding='utf-8') as f:
        f_csv = csv.writer(f)
        f_csv.writerow(headers)
        for t, c, u, p in zip(title, category, url, prod):
            f_csv.writerow([t, c, u, p])

save_csv()

这个版本的代码更健壮,既能正确处理带逗号的CSV字段,又能清理XML中的特殊字符,同时避免了空标签导致的报错。

内容的提问来源于stack exchange,提问作者Isaac Rivera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:08:24