You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫数据转DataFrame及CSV存储问题求助

解决requests_html爬取数据转DataFrame并导出CSV的问题

问题描述

使用requests_html库爬取某网站多页面的多链接数据,运行代码后可输出所需全部信息,但无法将数据转换为DataFrame并导出为CSV文件。怀疑需要将字典转换为列表,作为Python新手不知该如何推进,代码如下:

from requests_html import HTMLSession

s = HTMLSession()
def get_product_links(page):
  url = f'https://lakesshoweringspaces.com/catalogue-product-filter/page/{page}'
  links = []
  r = s.get(url)

  products = r.html.find("article.contentwrapper section.collection-wrapper-item")

  for item in products:
    links.append(item.find("a", first=True).attrs['href'].replace('?', ''))
  return links

#page1 = get_product_links(1)
#print(page1)

def parse_product(url):

  r = s.get(url)
  product_type = r.html.find('div.product-sidecontent h3', first=True).text.strip()
  collection = r.html.find('div.product-sidecontent h1', first=True).text.strip()
  description = r.html.find('div.information_grey_section h3.table-title', first=True).text.strip()
  detail = r.html.find('table', first=True)
  tabledata = [[c.text for c in row.find('td')] for row in detail.find('tr')][1:]
  tableheader = [[c.text for c in row.find('th')] for row in detail.find('tr')][0]
  table = [dict(zip(tableheader,t)) for t in tabledata]

  product ={
      'Product Type' : product_type,
      'Collection' : collection,
      'Short Description' : description,
      'Product Data' : table,
  }
  return product

results = []
for x in range(1, 10):
  print('Getting Page ', x)
  urls = get_product_links(x)
  for url in urls:
    print(parse_product(url))
    results.append(parse_product(url))
  print('Total Results: ', len(results))

问题原因

你的results列表中每个元素是包含嵌套列表的字典,其中Product Data字段是列表套字典的结构。而CSV是二维表格格式,无法直接存储这种嵌套的复杂数据结构,这就是无法正常导出的核心问题。需要先把嵌套数据扁平化,让每条数据对应CSV的一行。

修改后的代码

以下是调整后的代码,将产品基础信息与Product Data的每一行数据合并,生成扁平化条目,即可顺利转成DataFrame并导出CSV:

from requests_html import HTMLSession
import pandas as pd

s = HTMLSession()
def get_product_links(page):
    url = f'https://lakesshoweringspaces.com/catalogue-product-filter/page/{page}'
    links = []
    r = s.get(url)

    products = r.html.find("article.contentwrapper section.collection-wrapper-item")

    for item in products:
        links.append(item.find("a", first=True).attrs['href'].replace('?', ''))
    return links

def parse_product(url):
    r = s.get(url)
    product_type = r.html.find('div.product-sidecontent h3', first=True).text.strip()
    collection = r.html.find('div.product-sidecontent h1', first=True).text.strip()
    description = r.html.find('div.information_grey_section h3.table-title', first=True).text.strip()
    detail = r.html.find('table', first=True)
    tabledata = [[c.text for c in row.find('td')] for row in detail.find('tr')][1:]
    tableheader = [[c.text for c in row.find('th')] for row in detail.find('tr')][0]
    table = [dict(zip(tableheader,t)) for t in tabledata]

    # 扁平化数据:将基础信息与表格每行数据合并
    flattened_items = []
    for item in table:
        flattened_item = {
            'Product Type': product_type,
            'Collection': collection,
            'Short Description': description,
            **item  # 展开表格行的键值对
        }
        flattened_items.append(flattened_item)
    return flattened_items

results = []
for x in range(1, 10):
    print('Getting Page ', x)
    urls = get_product_links(x)
    for url in urls:
        product_items = parse_product(url)
        print(product_items)
        results.extend(product_items)  # 用extend添加列表中的每个元素
    print('Total Results: ', len(results))

# 转换为DataFrame并导出CSV
df = pd.DataFrame(results)
df.to_csv('shower_products.csv', index=False, encoding='utf-8-sig')
print("CSV文件已成功导出!")

关键修改说明

  1. 扁平化数据结构:修改parse_product函数,不再返回含嵌套列表的字典,而是生成多个扁平化字典,每个字典包含产品基础信息+表格单行数据,对应CSV的一行。
  2. 列表扩展方式:用results.extend(product_items)替代append,因为parse_product现在返回列表,extend会把列表里的每个元素单独加入results。
  3. DataFrame与CSV导出:导入pandas库,将results转为DataFrame,调用to_csv时设置index=False去掉索引列,encoding='utf-8-sig'避免中文乱码。

内容的提问来源于stack exchange,提问作者HRol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 12:35:31