Python爬虫数据转DataFrame及CSV存储问题求助
解决requests_html爬取数据转DataFrame并导出CSV的问题
问题描述
使用requests_html库爬取某网站多页面的多链接数据,运行代码后可输出所需全部信息,但无法将数据转换为DataFrame并导出为CSV文件。怀疑需要将字典转换为列表,作为Python新手不知该如何推进,代码如下:
from requests_html import HTMLSession s = HTMLSession() def get_product_links(page): url = f'https://lakesshoweringspaces.com/catalogue-product-filter/page/{page}' links = [] r = s.get(url) products = r.html.find("article.contentwrapper section.collection-wrapper-item") for item in products: links.append(item.find("a", first=True).attrs['href'].replace('?', '')) return links #page1 = get_product_links(1) #print(page1) def parse_product(url): r = s.get(url) product_type = r.html.find('div.product-sidecontent h3', first=True).text.strip() collection = r.html.find('div.product-sidecontent h1', first=True).text.strip() description = r.html.find('div.information_grey_section h3.table-title', first=True).text.strip() detail = r.html.find('table', first=True) tabledata = [[c.text for c in row.find('td')] for row in detail.find('tr')][1:] tableheader = [[c.text for c in row.find('th')] for row in detail.find('tr')][0] table = [dict(zip(tableheader,t)) for t in tabledata] product ={ 'Product Type' : product_type, 'Collection' : collection, 'Short Description' : description, 'Product Data' : table, } return product results = [] for x in range(1, 10): print('Getting Page ', x) urls = get_product_links(x) for url in urls: print(parse_product(url)) results.append(parse_product(url)) print('Total Results: ', len(results))
问题原因
你的results列表中每个元素是包含嵌套列表的字典,其中Product Data字段是列表套字典的结构。而CSV是二维表格格式,无法直接存储这种嵌套的复杂数据结构,这就是无法正常导出的核心问题。需要先把嵌套数据扁平化,让每条数据对应CSV的一行。
修改后的代码
以下是调整后的代码,将产品基础信息与Product Data的每一行数据合并,生成扁平化条目,即可顺利转成DataFrame并导出CSV:
from requests_html import HTMLSession import pandas as pd s = HTMLSession() def get_product_links(page): url = f'https://lakesshoweringspaces.com/catalogue-product-filter/page/{page}' links = [] r = s.get(url) products = r.html.find("article.contentwrapper section.collection-wrapper-item") for item in products: links.append(item.find("a", first=True).attrs['href'].replace('?', '')) return links def parse_product(url): r = s.get(url) product_type = r.html.find('div.product-sidecontent h3', first=True).text.strip() collection = r.html.find('div.product-sidecontent h1', first=True).text.strip() description = r.html.find('div.information_grey_section h3.table-title', first=True).text.strip() detail = r.html.find('table', first=True) tabledata = [[c.text for c in row.find('td')] for row in detail.find('tr')][1:] tableheader = [[c.text for c in row.find('th')] for row in detail.find('tr')][0] table = [dict(zip(tableheader,t)) for t in tabledata] # 扁平化数据:将基础信息与表格每行数据合并 flattened_items = [] for item in table: flattened_item = { 'Product Type': product_type, 'Collection': collection, 'Short Description': description, **item # 展开表格行的键值对 } flattened_items.append(flattened_item) return flattened_items results = [] for x in range(1, 10): print('Getting Page ', x) urls = get_product_links(x) for url in urls: product_items = parse_product(url) print(product_items) results.extend(product_items) # 用extend添加列表中的每个元素 print('Total Results: ', len(results)) # 转换为DataFrame并导出CSV df = pd.DataFrame(results) df.to_csv('shower_products.csv', index=False, encoding='utf-8-sig') print("CSV文件已成功导出!")
关键修改说明
- 扁平化数据结构:修改
parse_product函数,不再返回含嵌套列表的字典,而是生成多个扁平化字典,每个字典包含产品基础信息+表格单行数据,对应CSV的一行。 - 列表扩展方式:用
results.extend(product_items)替代append,因为parse_product现在返回列表,extend会把列表里的每个元素单独加入results。 - DataFrame与CSV导出:导入
pandas库,将results转为DataFrame,调用to_csv时设置index=False去掉索引列,encoding='utf-8-sig'避免中文乱码。
内容的提问来源于stack exchange,提问作者HRol
相关产品推荐
相关产品推荐

