Python移除XML特殊字符及CSV导出逗号分隔问题求助
解决CSV字段拆分与XML特殊字符清理问题
咱们一步步来解决你的两个问题:先是CSV里逗号导致内容拆分的问题,然后是XML特殊字符的清理。
一、修复CSV中逗号导致的内容拆分问题
你现在遇到的CSV拆分问题,根源是手动拼接字符串写入文件——这种方式完全绕开了csv模块的内置处理逻辑,自然会让带逗号的内容被错误拆分。
csv模块本身就专门处理这种场景:当字段包含逗号、引号或者换行符时,它会自动用双引号包裹该字段,确保内容完整保留。只需要修改你的save_csv函数,使用writerow()方法逐行写入即可:
def save_csv(): # 添加newline=''是csv模块的推荐写法,避免跨平台换行符混乱 with open('titles.csv', 'w', newline='', encoding='utf-8') as f: f_csv = csv.writer(f) f_csv.writerow(headers) # 遍历配对后的列表,用writerow写入每一行 for t, c, u, p in zip(title, category, url, prod): f_csv.writerow([t, c, u, p])
这样修改后,像Add, Change, or Remove Transitions between Slides这类带逗号的标题,会被自动包裹在双引号里写入CSV,打开时就不会被拆分成多个单元格了。
如果你确实有需求要移除标题中的逗号(不推荐,会破坏原始内容),可以在提取标题时用replace方法:
title.append(t.find('title').text.replace(',', ''))
但显然,用csv模块的原生方法保留原内容是更合理的方案。
二、清理XML中的特殊字符
XML里的特殊字符主要分两种情况:
- XML实体引用:比如
&(对应&)、<(对应<)、>(对应>)等,xml.etree.ElementTree在解析XML时会自动把这些实体转换成对应的正常字符,不需要额外处理。 - 未转义的无效字符:比如XML中直接出现
<(而不是<),或者一些非打印控制字符,这类情况会导致解析报错,或者提取出的文本有残留垃圾字符。
我们可以写一个清理函数来处理这些情况:
import re def clean_special_chars(text): if text is None: return '' # 替换XML实体引用为对应正常字符 text = re.sub(r'&', '&', text) text = re.sub(r'<', '<', text) text = re.sub(r'>', '>', text) text = re.sub(r'"', '"', text) text = re.sub(r''', "'", text) # 移除XML不允许的控制字符(可选,根据你的需求调整) text = re.sub(r'[\x00-\x08\x0B\x0C\x0E-\x1F]', '', text) return text
然后在提取文本的时候调用这个函数,比如:
title.append(clean_special_chars(t.find('title').text))
整合后的完整代码
我还优化了你的数据提取逻辑,减少了重复遍历XML节点的次数,同时添加了空值判断,避免某个标签不存在时抛出错误:
import xml.etree.ElementTree as ET import csv import re def clean_special_chars(text): if text is None: return '' # 解析XML实体引用 text = re.sub(r'&', '&', text) text = re.sub(r'<', '<', text) text = re.sub(r'>', '>', text) text = re.sub(r'"', '"', text) text = re.sub(r''', "'", text) # 移除无效控制字符 text = re.sub(r'[\x00-\x08\x0B\x0C\x0E-\x1F]', '', text) return text tree = ET.parse('file.xml') root = tree.getroot() title = [] category = [] url = [] prod = [] def extract_data(): # 一次遍历所有solution节点,避免重复遍历 for solution in root.findall('solution'): # 处理head中的title head = solution.find('head') title_elem = head.find('title') if head is not None else None title.append(clean_special_chars(title_elem.text if title_elem is not None else '')) # 处理body中的各个字段 body = solution.find('body') if body is not None: cat_elem = body.find('category') category.append(clean_special_chars(cat_elem.text if cat_elem is not None else '')) video_elem = body.find('video') url.append(clean_special_chars(video_elem.text if video_elem is not None else '')) prod_elem = body.find('product') prod.append(clean_special_chars(prod_elem.text if prod_elem is not None else '')) else: # 如果body不存在,添加空字符串占位 category.append('') url.append('') prod.append('') extract_data() headers = ['Title', 'Category', 'Video URL','Product'] def save_csv(): with open('titles.csv', 'w', newline='', encoding='utf-8') as f: f_csv = csv.writer(f) f_csv.writerow(headers) for t, c, u, p in zip(title, category, url, prod): f_csv.writerow([t, c, u, p]) save_csv()
这个版本的代码更健壮,既能正确处理带逗号的CSV字段,又能清理XML中的特殊字符,同时避免了空标签导致的报错。
内容的提问来源于stack exchange,提问作者Isaac Rivera
相关产品推荐
相关产品推荐

