You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7+BeautifulSoup:如何拆分提取文本至CSV对应列?

解决Python 2.7 + BeautifulSoup爬取数据写入CSV的列对齐问题

看来你在把爬取到的球员数据写入CSV时,遇到了表头和内容都挤在同一列的麻烦——这通常是因为你把整个表头元组、整个球员信息都当成单个元素写入了,CSV自然会把它们塞进同一个单元格里。下面我给你一套完整的解决方案,帮你把每个字段拆到单独列里:

核心思路

  1. 先明确要输出的表头字段,把它们拆成一个字符串列表(不是元组套列表),这样CSV写入器会自动把每个字段放到单独列。
  2. 从页面中提取球员的每个字段(Name、Position等),把它们和ImageURL一起组成一个列表/字典,确保每个字段对应表头的一列。
  3. 使用Python2.7的csv模块正确写入,避免编码和格式问题。

完整代码示例

1. 导入依赖模块

import csv
from bs4 import BeautifulSoup
import urllib2

2. 爬取页面并提取数据

# 替换成你的目标URL
target_url = "https://example.com/players"
response = urllib2.urlopen(target_url)
# 用html.parser解析,如果你装了lxml也可以换成'lxml'
soup = BeautifulSoup(response.read(), 'html.parser')

# 假设页面中每个球员的信息都包裹在class为'player-card'的div里,根据实际页面调整
player_cards = soup.find_all('div', class_='player-card')

3. 写入CSV文件

这里提供两种写法,选你觉得顺手的就行:

方法一:使用csv.writer(基础写法)

# 定义表头列表,每个元素对应一列
headers = ['Name', 'Position', 'Height', 'Weight', 'Born', 'ImageURL']

# Python2.7写入CSV必须用'wb'模式,避免换行符乱码
with open('players_data.csv', 'wb') as csv_file:
    writer = csv.writer(csv_file)
    # 先写入表头,每个字段自动分到单独列
    writer.writerow(headers)
    
    for card in player_cards:
        # 提取每个字段,注意:这里的选择器要和你爬取的页面HTML结构完全匹配!
        # 比如页面中球员名字在h3标签,class是'player-name',就这么写
        name = card.find('h3', class_='player-name').text.strip().encode('utf-8')
        position = card.find('span', class_='player-position').text.strip().encode('utf-8')
        height = card.find('span', class_='player-height').text.strip().encode('utf-8')
        weight = card.find('span', class_='player-weight').text.strip().encode('utf-8')
        born = card.find('span', class_='player-born').text.strip().encode('utf-8')
        image_url = card.find('img')['src'].encode('utf-8')
        
        # 把所有字段组成一个列表,每个元素对应表头的一列
        player_row = [name, position, height, weight, born, image_url]
        writer.writerow(player_row)

方法二:使用csv.DictWriter(更清晰,不易出错)

这种方法用字典对应字段名,不用担心顺序写错:

headers = ['Name', 'Position', 'Height', 'Weight', 'Born', 'ImageURL']

with open('players_data.csv', 'wb') as csv_file:
    writer = csv.DictWriter(csv_file, fieldnames=headers)
    # 写入表头
    writer.writeheader()
    
    for card in player_cards:
        # 用字典存储每个球员的字段,键对应表头
        player_dict = {
            'Name': card.find('h3', class_='player-name').text.strip().encode('utf-8'),
            'Position': card.find('span', class_='player-position').text.strip().encode('utf-8'),
            'Height': card.find('span', class_='player-height').text.strip().encode('utf-8'),
            'Weight': card.find('span', class_='player-weight').text.strip().encode('utf-8'),
            'Born': card.find('span', class_='player-born').text.strip().encode('utf-8'),
            'ImageURL': card.find('img')['src'].encode('utf-8')
        }
        writer.writerow(player_dict)

关键注意事项

  • 编码处理:Python2.7默认字符串是ASCII,所以提取到非ASCII字符(比如中文、特殊符号)时,一定要用.encode('utf-8')转码,否则写入CSV会报错或者乱码。
  • 选择器匹配:代码里的find参数(标签名、class_)必须和你爬取的页面HTML结构完全一致,比如页面中球员位置在div而不是span,就要改成card.find('div', class_='player-position')。
  • CSV模式:Python2.7写入CSV必须用'wb'(二进制写入)模式,否则Windows下会出现多余的空行。

这样处理后,你的CSV文件里每个表头都会单独占一列,球员的对应信息也会准确对应到各自的列里,ImageURL也会单独成一列啦!

内容的提问来源于stack exchange,提问作者Bab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:42:57