You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:如何拆分长列表数据适配Dataframe?

问题描述

我是编程新手,正在制作爬取BoardGameGeek账户的网页爬虫,尝试抓取table并转换为Dataframe。已获取表头列表,但表格数据为长列表,共8个表头却有483个数据项,无法适配Dataframe。请问如何拆分该长列表数据以适配Dataframe?

原代码:

# Import libraries
import requests
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np

# Create an URL object
url = 'https://boardgamegeek.com/collection/user/kyletravels?want=1&subtype=boardgame&ff=1'
# Create object page
pages = requests.get(url)
pages.text

# parser-lxml = Change html to Python friendly format
soup = BeautifulSoup(pages.text, 'lxml')

#Access <tbody> tag
table = soup.table

# Obtain information from tag <table>
table1 = soup.find('table', id='collectionitems')
table1

# Obtain every title of columns with tag <th>
headers = []
for i in table1.find_all('th'):
 title = i.text
 headers.append(title)

info = []
for i in table1.find_all('td'):
    stats = i.text
    info.append(stats)


total = pd.DataFrame(data=headers, columns=info)
问题分析

你现在的核心问题有两个:

  1. 所有表格单元格内容被塞进了一个一维列表info,但DataFrame需要二维结构(每行对应一组和表头数量匹配的单元格)。
  2. 最后创建DataFrame的写法完全颠倒,把表头当成数据、数据当成列名,逻辑错误。
解决方案

把一维的info列表按表头数量(8个)拆分成二维列表,每8个元素对应表格的一行,再用这个二维列表作为DataFrame的数据,表头作为列名即可。

修正后的代码:

# Import libraries
import requests
from bs4 import BeautifulSoup
import pandas as pd
import numpy as np

# Create an URL object
url = 'https://boardgamegeek.com/collection/user/kyletravels?want=1&subtype=boardgame&ff=1'
# Create object page
pages = requests.get(url)

# parser-lxml = Change html to Python friendly format
soup = BeautifulSoup(pages.text, 'lxml')

# Obtain information from tag <table>
table1 = soup.find('table', id='collectionitems')

# Obtain every title of columns with tag <th>
headers = []
for i in table1.find_all('th'):
    title = i.text.strip()  # 清理表头里的换行、空格
    headers.append(title)

# 获取所有表格单元格内容
info = []
for i in table1.find_all('td'):
    stats = i.text.strip()  # 清理单元格文本冗余内容
    info.append(stats)

# 按表头数量拆分列表,每8个元素为一行
col_count = len(headers)
table_data = [info[i:i+col_count] for i in range(0, len(info), col_count)]

# 创建正确的DataFrame
total = pd.DataFrame(data=table_data, columns=headers)

# 查看前5行结果
print(total.head())
关键说明
  • strip():清理文本中的换行、空格等冗余字符,让数据更整洁。
  • 列表拆分:通过range(0, len(info), col_count)按步长切割列表,自动将一维列表转为每行8个元素的二维结构,完美匹配表头数量。
  • DataFrame参数:data传入二维数据,columns传入表头列表,修正了你原代码的参数颠倒问题。

内容的提问来源于stack exchange,提问作者KB53

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 09:09:18