Python网页爬取:如何拆分长列表数据适配Dataframe?
问题描述
我是编程新手,正在制作爬取BoardGameGeek账户的网页爬虫,尝试抓取table并转换为Dataframe。已获取表头列表,但表格数据为长列表,共8个表头却有483个数据项,无法适配Dataframe。请问如何拆分该长列表数据以适配Dataframe?
原代码:
# Import libraries import requests from bs4 import BeautifulSoup import pandas as pd import numpy as np # Create an URL object url = 'https://boardgamegeek.com/collection/user/kyletravels?want=1&subtype=boardgame&ff=1' # Create object page pages = requests.get(url) pages.text # parser-lxml = Change html to Python friendly format soup = BeautifulSoup(pages.text, 'lxml') #Access <tbody> tag table = soup.table # Obtain information from tag <table> table1 = soup.find('table', id='collectionitems') table1 # Obtain every title of columns with tag <th> headers = [] for i in table1.find_all('th'): title = i.text headers.append(title) info = [] for i in table1.find_all('td'): stats = i.text info.append(stats) total = pd.DataFrame(data=headers, columns=info)
问题分析
你现在的核心问题有两个:
- 所有表格单元格内容被塞进了一个一维列表
info,但DataFrame需要二维结构(每行对应一组和表头数量匹配的单元格)。 - 最后创建DataFrame的写法完全颠倒,把表头当成数据、数据当成列名,逻辑错误。
解决方案
把一维的info列表按表头数量(8个)拆分成二维列表,每8个元素对应表格的一行,再用这个二维列表作为DataFrame的数据,表头作为列名即可。
修正后的代码:
# Import libraries import requests from bs4 import BeautifulSoup import pandas as pd import numpy as np # Create an URL object url = 'https://boardgamegeek.com/collection/user/kyletravels?want=1&subtype=boardgame&ff=1' # Create object page pages = requests.get(url) # parser-lxml = Change html to Python friendly format soup = BeautifulSoup(pages.text, 'lxml') # Obtain information from tag <table> table1 = soup.find('table', id='collectionitems') # Obtain every title of columns with tag <th> headers = [] for i in table1.find_all('th'): title = i.text.strip() # 清理表头里的换行、空格 headers.append(title) # 获取所有表格单元格内容 info = [] for i in table1.find_all('td'): stats = i.text.strip() # 清理单元格文本冗余内容 info.append(stats) # 按表头数量拆分列表,每8个元素为一行 col_count = len(headers) table_data = [info[i:i+col_count] for i in range(0, len(info), col_count)] # 创建正确的DataFrame total = pd.DataFrame(data=table_data, columns=headers) # 查看前5行结果 print(total.head())
关键说明
strip():清理文本中的换行、空格等冗余字符,让数据更整洁。- 列表拆分:通过
range(0, len(info), col_count)按步长切割列表,自动将一维列表转为每行8个元素的二维结构,完美匹配表头数量。 - DataFrame参数:
data传入二维数据,columns传入表头列表,修正了你原代码的参数颠倒问题。
内容的提问来源于stack exchange,提问作者KB53
相关产品推荐
相关产品推荐

