You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Beautiful Soup爬虫结果转为水平结构Pandas DataFrame并优化代码

嘿,别担心,新手遇到这些问题太正常了!我来一步步帮你解决这三个核心需求:把抓取的数据转成结构一致的DataFrame、重构代码为Python风格的函数,还有避开嵌套循环的坑。

1. 先把代码重构为Python风格的函数

新手的代码通常是线性的,把重复的逻辑(请求网页、解析、数据提取)封装成函数会更清晰、复用性强,也符合Python的简洁风格。这里给你一个通用的重构示例,你可以根据自己的目标页面调整选择器:

from bs4 import BeautifulSoup
import requests
import pandas as pd

def scrape_target_website(url):
    # 1. 网页请求(加基础异常处理,避免请求崩溃)
    try:
        response = requests.get(url)
        response.raise_for_status()  # 自动抛出请求错误(比如404、500)
    except requests.exceptions.RequestException as e:
        print(f"请求网页出错啦: {e}")
        return []
    
    # 2. 解析HTML内容
    soup = BeautifulSoup(response.text, 'html.parser')
    
    # 3. 提取数据(这里需要你替换成目标页面的实际选择器)
    # 假设你要抓取页面上的所有商品条目
    item_containers = soup.find_all('div', class_='product-card')  # 替换成你页面的条目容器选择器
    scraped_data = []
    
    for container in item_containers:
        # 提取单个条目的所有字段,用字典存储
        item_info = {
            '商品名称': container.find('h3', class_='product-title').get_text(strip=True) if container.find('h3', class_='product-title') else '无数据',
            '商品价格': container.find('span', class_='price-tag').get_text(strip=True) if container.find('span', class_='price-tag') else '无数据',
            '详情链接': container.find('a')['href'] if container.find('a') else '无链接'
        }
        scraped_data.append(item_info)
    
    return scraped_data

这个函数的好处是:逻辑模块化,输入是目标URL,输出是包含所有条目数据的列表,后续调试和修改都很方便。

2. 生成和print(data)结构一致的DataFrame

你提到现有代码展示的是垂直排列的一行数据,应该是之前要么只抓取了单个条目,要么是逐个打印字典的键值对。要让DataFrame呈现水平结构(每列对应一个字段,每行对应一条数据),只需要把函数返回的列表直接传给pd.DataFrame()即可:

# 调用函数抓取数据
target_url = "你的目标网页URL"
data_list = scrape_target_website(target_url)

# 转成DataFrame
df = pd.DataFrame(data_list)

# 查看结果,和你想要的水平结构完全一致
print(df)

如果你的data是单个字典(比如只抓取了一条数据),记得用列表包裹后再转DataFrame:pd.DataFrame([data]),这样就会生成一行多列的水平结构,而不是一列多行的垂直结构。

3. 避开嵌套循环+append/concat的坑

你之前用append/concat报错,大概率是因为在循环里一次次给DataFrame添加单行数据——这种做法不仅效率极低,还容易因为数据结构不统一引发错误。Python的最佳实践是:先把所有数据收集到一个列表里(每个元素是一个字典,对应一条数据的所有字段),最后一次性转成DataFrame。

比如你之前可能写了类似这样的错误代码:

# 错误示例:循环append,低效且易报错
df = pd.DataFrame()
for item in item_containers:
    title = item.find('h3').text
    price = item.find('span', class_='price').text
    df = df.append({'title': title, 'price': price}, ignore_index=True)

而正确的做法就是我上面函数里的逻辑:先把所有条目字典存入scraped_data列表,最后用pd.DataFrame(scraped_data)一次性生成DataFrame——既不会报错,运行效率也提升很多。

完整示例串起来

最后给你一个可以直接运行的完整代码,你只需要替换目标URL和页面选择器:

from bs4 import BeautifulSoup
import requests
import pandas as pd

def scrape_target_website(url):
    try:
        response = requests.get(url)
        response.raise_for_status()
    except requests.exceptions.RequestException as e:
        print(f"请求失败: {e}")
        return []
    
    soup = BeautifulSoup(response.text, 'html.parser')
    item_containers = soup.find_all('div', class_='product-card')  # 替换成你的条目容器选择器
    
    scraped_data = []
    for container in item_containers:
        item_info = {
            '商品名称': container.find('h3', class_='product-title').get_text(strip=True) if container.find('h3', class_='product-title') else '无数据',
            '商品价格': container.find('span', class_='price-tag').get_text(strip=True) if container.find('span', class_='price-tag') else '无数据',
            '评分': container.find('span', class_='rating-star').get_text(strip=True) if container.find('span', class_='rating-star') else '无数据'
        }
        scraped_data.append(item_info)
    
    return scraped_data

# 主程序入口
if __name__ == "__main__":
    target_url = "https://example.com/products"  # 替换成你的目标URL
    data = scrape_target_website(target_url)
    
    if data:
        df = pd.DataFrame(data)
        print("抓取的DataFrame结果:")
        print(df)
        # 可选:保存到CSV文件
        df.to_csv('scraped_products.csv', index=False, encoding='utf-8-sig')
    else:
        print("没有抓取到任何数据哦")

内容的提问来源于stack exchange,提问作者Margosia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:57:21