Python使用BeautifulSoup导出CSV时内容出现b或\n前缀问题求解
问题排查与修正方案
核心问题1:b前缀出现的原因
你在写入CSV时对n/p/l三个变量调用了.encode('utf-8'),该方法会将字符串转为bytes字节类型,CSV模块写入非字符串类型时,会直接把字节对象的字面表示(带b''包裹)写入文件,就出现了你看到的前缀。
核心问题2:多余换行/空行出现的原因
- 你调用
name.get_text('title')的参数使用错误:get_text()第一个参数是用来拼接多个子节点文本的分隔符,你传'title'完全没必要,直接用.get_text().strip()就能清除文本前后的换行、空格等空白字符。 - 打开CSV文件时没有指定
newline=''参数,Windows系统下会自动额外插入换行符,导致每行之间多空行。 - 你最后写的
file.close没有加括号,属于只写了方法名没有实际调用关闭操作,文件缓存可能没有正常写入。
修正后的完整代码
from bs4 import BeautifulSoup import requests import csv source = requests.get('https://www.shopdisney.com/uniquely-disney/parks-inspired/walt-disney-world-50th-collection/').text soup = BeautifulSoup(source, 'html.parser') # 改用with上下文管理器,自动处理文件关闭,指定newline避免空行,指定编码兼容特殊字符 with open('Disney50th.csv', 'w', newline='', encoding='utf-8-sig') as file: writer = csv.writer(file) # 写入表头 writer.writerow(['Item Name', 'Item Price', 'Item Link']) products = soup.find_all('div', class_="product__tile") for item_info in products: price = item_info.find('span', class_='value') name = item_info.find('a', class_="product__tile_link") link = item_info.find('a', class_="product__tile_link") p = price.attrs['content'] n = name.get_text().strip() # 清除前后空白、换行 l = link.attrs['href'] print("Item Name =", n, '\n' "Item Link =", l, '\n' "Item Price =",p) # 直接写入字符串,不需要encode writer.writerow([n, p, l])
补充说明
- 打开文件时指定
encoding='utf-8-sig'是为了让Windows下用Excel打开CSV时不会出现中文乱码,如果你不需要用Excel打开可以换成utf-8。 - 用with上下文管理器就不需要手动写close方法,程序会自动在代码块结束后关闭文件,避免资源泄露。
内容的提问来源于stack exchange,提问作者Omece
相关产品推荐
相关产品推荐

