You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取结果写入Excel时,如何按列拆分含逗号的数据?

问题描述

我知道常规逗号分隔的数据(比如 x, y, z)怎么按列拆分,但遇到字段本身包含逗号的情况就不知道该处理了,比如这类格式:

Doe, John, BookName, Year, abstract with commas, links.

爬取后的数据在Excel里显示错乱,我希望每个字段能正确对应到目标列中。

我的代码
from unittest import result
import requests
from bs4 import BeautifulSoup
import csv
import urllib3.request
urllib3.disable_warnings(urllib3.exceptions.InsecureRequestWarning)


fakdep = '165'  
offset = input('Please enter number of offset:')
url = 'https://repositori.usu.ac.id/handle/123456789/{}?offset={}'.format(fakdep,offset)
headers = {
    "user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/105.0.0.0 Safari/537.36"
    }

datas = []
count_page = 0
for page in range(1,2): 
    count_page+=1
    print('Scraping Offset No:', count_page)
    result = requests.get(url+str(page), verify=False) 
    
    soup = BeautifulSoup(result.text, 'html.parser')
    items = soup.find_all('li','ds-artifact-item')
    for it in items:
        author = it.find('span','author h4').text
        title = ''.join(it.find('a',href=True).text.strip().split('\n'))
        year = it.find('span','date').text
        abstract = ''.join(it.find('div','artifact-abstract').text.strip().split('\n'))
        link = it.find('a')['href']
        datas.append([author, title, year, abstract, link])

kepala = ['Author', 'Title', 'Year', 'Abstract', 'Link']
thewriter = csv.writer(open('results/{}_{}.csv'.format(fakdep,offset), 'w', newline=''))
thewriter.writerow(kepala)
for d in datas: thewriter.writerow(d)
解决方法

问题根源是CSV写入时,字段内部的逗号被Excel当成了列分隔符。Python内置的csv模块已经有现成的处理方案,只需在创建csv.writer时指定quoting参数,让包含特殊字符的字段自动被引号包裹,避免解析错误。

修改代码中的CSV写入部分:

# 替换原有的thewriter创建语句,选择以下两种方式之一即可
# 方式1:所有字段都用引号包裹,最稳妥
thewriter = csv.writer(open('results/{}_{}.csv'.format(fakdep,offset), 'w', newline='', encoding='utf-8'), quoting=csv.QUOTE_ALL)

# 方式2:仅当字段包含逗号、引号或换行时才加引号,更简洁
# thewriter = csv.writer(open('results/{}_{}.csv'.format(fakdep,offset), 'w', newline='', encoding='utf-8'), quoting=csv.QUOTE_MINIMAL)

同时建议添加encoding='utf-8'参数,避免导出的CSV出现中文乱码问题。修改后生成的文件,Excel打开时会正确识别每个字段,不会因内部逗号拆分列。

内容的提问来源于stack exchange,提问作者Milkwiway

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 00:25:28