You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量转置Pandas读取的多DataFrame并合并为单表?

Efficient Batch Processing for Transposing and Merging Scraped Tables in Pandas

Hey there! I’ve run into similar headaches dealing with scraped tables in Pandas, so let’s break down how to fix this efficiently and cleanly.

The Core Issue with Your Current Method

Your original approach concatenates all raw tables first, then tries to slap on a universal header/value structure. This fails because each table’s first row is actually its own header—so you end up mixing duplicate headers into your data before transposing, which messes up the final format. Plus, row-by-row appends are way slower than batch operations.

Step-by-Step Efficient Solution

We’ll process each table individually first to turn it into a single row, then merge all rows together. This ensures column alignment and avoids duplicate headers, while keeping operations fast.

  1. Convert each table to a single row
    For every scraped table, turn its header-value pairs into a single row where each header becomes a column name, and the corresponding value fills the cell.
  2. Collect processed rows
    Gather these single-row DataFrames in a list (this is way faster than appending one-by-one to a main DataFrame).
  3. Merge with auto-alignment
    Use pd.concat() to combine all rows—Pandas will automatically match columns across tables, filling missing ones with NaN (empty values) for tables that don’t have those headers.

Full Code Implementation

import pandas as pd
from bs4 import BeautifulSoup

# Assume you already have your soup object from scraping
tables = soup.find_all('table')

# Initialize a list to hold our processed single-row DataFrames
processed_rows = []

for table in tables:
    # Read the individual table into a DataFrame
    df = pd.read_html(str(table))[0]
    
    # Label columns for header-value pair structure
    df.columns = ['header', 'value']
    
    # Transpose to turn rows into columns, then reset index to get a clean single row
    transposed_row = df.set_index('header').T.reset_index(drop=True)
    
    # Add the processed row to our list
    processed_rows.append(transposed_row)

# Merge all processed rows into one final DataFrame
final_df = pd.concat(processed_rows, ignore_index=True)

Key Improvements Explained

  • No duplicate headers: By processing each table before merging, we ensure each original header becomes a unique column in the final output.
  • Auto-fill missing columns: Pandas’ concat() handles column alignment automatically—if one table has 9 columns and another has 12, the missing 3 columns will be filled with NaN for the shorter table’s row.
  • Faster performance: Collecting DataFrames in a list and concatenating once is drastically more efficient than appending rows individually (Pandas is optimized for batch operations like this).

Bonus Optimization

If you want to skip converting each table object to a string, you can pass the prettified BeautifulSoup element directly to read_html (works in most recent Pandas versions):

df = pd.read_html(table.prettify())[0]

内容的提问来源于stack exchange,提问作者Chris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:02:12