如何合并23个同格式HTML文件生成统一的Pandas DataFrame?
Solution to Merge HTML Tables into a Single DataFrame
Let's break down what's going wrong and how to fix it:
The Problem in Your Current Code
When you use pd.read_html(filename), it returns a list of DataFrames (even if your HTML only contains one table). Right now, you're appending this entire list to list_data, which means list_data ends up being a list of lists (each entry is a list containing one DataFrame). That's why pd.concat(list_data) isn't producing the merged DataFrame you want.
Corrected Code
First, adjust your loop to extract the actual DataFrame from the list returned by pd.read_html:
import glob import pandas as pd # Get all HTML files (use raw string to avoid escape character issues) all_rec = glob.glob(r'D:\python\*.html') list_data = [] for filename in all_rec: # Read HTML tables - returns a list of DataFrames tables = pd.read_html(filename) # Append the first (and only) DataFrame from each file to list_data list_data.append(tables[0]) # Merge all DataFrames into one single DataFrame merged_df = pd.concat(list_data, ignore_index=True)
Explanation
tables[0]: Since each of your HTML files has the same single table structure,tables[0]grabs the first (and only) DataFrame from the list returned bypd.read_html. This ensureslist_databecomes a flat list of DataFrames instead of nested lists.pd.concat(list_data, ignore_index=True): Now thatlist_datacontains direct references to each DataFrame, passing it topd.concatwill stack all rows together into one big DataFrame. Theignore_index=Trueparameter resets the index to be continuous across all merged data.
Verify the Result
You can confirm the merged DataFrame is correct by checking its dimensions:
print(f"Total rows in merged DataFrame: {len(merged_df)}") print(f"Total columns: {len(merged_df.columns)}")
内容的提问来源于stack exchange,提问作者jatin
相关产品推荐
相关产品推荐

