创建Pandas DataFrame报错All arrays must be of the same length求解决方案
问题排查与修复:Pandas创建DataFrame时报错“All arrays must be of the same length”
错误原因分析
你的代码存在两个核心问题导致长度不匹配:
- 循环内全局查询重复添加数据:在遍历每条
quote时,你用doc.find_all在整个文档中查询所有quote、作者和标签,这会导致每次循环都把页面上的10条数据全部重复添加到Quotes和Quotes_by列表中。最终这两个列表的长度是10*10=100,而Tags列表只在每次循环添加1个元素,长度为10,三者长度不一致,触发Pandas的报错。 - Tags列表结构不匹配:当前代码往
Tags里添加的是包含quote和tags的字典,而另外两个列表是纯文本列表,结构不统一,进一步加剧了数据长度和格式的问题。
修复后的代码
import requests from bs4 import BeautifulSoup import pandas as pd # 连接目标URL quotes_url = "https://quotes.toscrape.com/" response = requests.get(quotes_url) Quotes = [] Quotes_by = [] Tags = [] # 检查请求状态 if response.status_code == 200: url_contents = response.text doc = BeautifulSoup(url_contents, 'html.parser') quote_rows = doc.find_all('div', {'class': 'quote'}) # 遍历每条quote元素,仅从当前元素内提取数据 for quote in quote_rows: # 提取当前quote的文本 quote_text = quote.find('span', {'class': 'text'}).text.strip() Quotes.append(quote_text) # 提取当前quote的作者 author = quote.find('small', {'class': 'author'}).text.strip() Quotes_by.append(author) # 提取当前quote的所有标签,整理为列表 tag_elements = quote.find_all('a', {'class': 'tag'}) tag_list = [tag.text.strip() for tag in tag_elements] Tags.append(tag_list) # 构造字典并创建DataFrame quotes_dict = { 'Author': Quotes_by, 'Tags': Tags, 'Quote': Quotes } Quotes_df = pd.DataFrame(quotes_dict) print(Quotes_df) Quotes_df.to_csv('quote.csv', index=None) else: print(f"Error: {response.status_code}")
关键修改点
- 局部数据提取:将
doc.find_all改为quote.find/quote.find_all,确保每次只从当前遍历的quote元素中提取对应的数据,避免重复添加。 - 统一列表结构:
Tags列表现在存储的是每条quote对应的标签列表,和Quotes、Quotes_by保持相同长度(均为10),且格式匹配。 - 简化逻辑:去掉冗余的嵌套循环,直接在单次遍历中完成数据提取和列表追加,代码更简洁高效。
内容的提问来源于stack exchange,提问作者Terryktee
相关产品推荐
相关产品推荐

