You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping爬取游戏站点导出CSV出现多余列问题求助

问题原因

出现额外列是两个问题共同导致的:

  • 写入表头时末尾没有添加换行符\n,第一条数据会直接拼接在表头行的末尾,造成字段整体错位
  • 手动用逗号拼接CSV字段,没有对字段内容做转义处理。如果爬取到的游戏名、数值等字段本身包含逗号,会被pandas读取时识别为列分隔符,把单个字段拆分为多个列,就会出现你看到的['Assassins Creed']、['7.6']这类多余列

修复方案

推荐使用Python内置的csv模块完成CSV写入,模块会自动处理特殊字符转义、字段包裹,避免手动拼接的问题,修改后的参考代码如下:

import csv
import requests
from bs4 import BeautifulSoup
import pandas as pd

# 写入CSV部分修改为csv模块实现
with open('urls.txt', 'r') as inf:
    # 定义表头
    headers = ['Name', 'Note', 'Critique', 'Coup_de_coeur', 'Envie', 'Joueur']
    with open('ubisoft1.csv','w', encoding='utf8', newline='') as f:
        writer = csv.writer(f)
        # 写入表头
        writer.writerow(headers)
        for row in inf:
            url = row.strip()
            response = requests.get(url)
            if response.ok:
                soup = BeautifulSoup(response.text, 'html.parser')
                # 注意class属性不要拆成两行写,会匹配失败
                title = soup.find('h1', attrs={'class': 'pvi-product-title'}).text.strip()
                note_redac = soup.find('span',attrs={'class': 'pvi-scrating-value'}).text.strip()
                stats = soup.select('b',attrs={'class': 'pvi-stats-number '})
                nb_critique = stats[0].text.strip()
                coup_de_coeur = stats[1].text.strip()
                envie_de_jouer = stats[2].text.strip()
                joueur_ponctuelle = stats[3].text.strip()
                # 直接传列表给writerow,自动处理拼接和转义
                writer.writerow([title, note_redac, nb_critique, coup_de_coeur, envie_de_jouer, joueur_ponctuelle])

df = pd.read_csv('ubisoft1.csv')
df.head()

内容的提问来源于stack exchange,提问作者Nitsau

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 00:42:03