韩国美妆网站爬虫Excel写入异常:Pandas存问题xlrd正常求解决方案
description字段 Hey there! Let's break down why you're seeing that unwanted description column when using Pandas to write Excel files, and how to fix it while sticking with Pandas (which is totally the right call for convenience).
核心原因
The key difference between your xlrd workflow and Pandas workflow is how each handles column selection:
- When using xlrd, you were probably manually specifying exactly which columns to write (e.g., looping through specific fields and writing cell by cell). This means you naturally excluded the
descriptionfield without even thinking about it. - Pandas, by contrast, writes all columns present in your DataFrame by default. So if the
descriptionfield is lurking in your DataFrame (even if you didn't notice it), Pandas will include it in the Excel output.
可行解决方案
Here are a few straightforward fixes to keep using Pandas while ditching that unwanted column:
1. 先确认DataFrame的列结构
First, verify if description is actually in your DataFrame. Run these lines before writing to Excel:
print("Current DataFrame columns:", df.columns.tolist()) # Or get a detailed overview df.info()
This will confirm if the column exists—9 times out of 10, it does, and that's the root issue.
2. 明确指定要写入的列(推荐)
Instead of letting Pandas write all columns, explicitly list the ones you want to include. This is the most maintainable approach because it makes your intent clear:
# Replace with your actual desired column names desired_columns = ['product_name', 'price', 'brand', 'category'] df[desired_columns].to_excel('korean_beauty_data.xlsx', index=False)
The index=False parameter ensures you don't write Pandas' default row index to the Excel file.
3. 删除不需要的列再写入
If you prefer to modify the DataFrame directly, drop the description column before writing:
# Drop the column (inplace=True modifies the DataFrame directly) df.drop(columns=['description'], inplace=True) # Now write to Excel df.to_excel('korean_beauty_data.xlsx', index=False)
If you don't want to modify the original DataFrame, skip inplace=True and assign to a new variable:
filtered_df = df.drop(columns=['description']) filtered_df.to_excel('korean_beauty_data.xlsx', index=False)
4. 在爬虫数据阶段提前过滤
To prevent the description field from ever entering your DataFrame, filter it out when processing the raw scraped data. For example, if your crawler returns a list of dictionaries:
# raw_data is the list of dictionaries from your crawler filtered_raw_data = [ {key: value for key, value in item.items() if key != 'description'} for item in raw_data ] # Now convert to DataFrame df = pd.DataFrame(filtered_raw_data)
This way, the description field never makes it into your DataFrame in the first place.
总结
The issue boils down to Pandas writing all DataFrame columns by default, whereas your xlrd workflow was more selective. Any of the above fixes will let you keep using Pandas' convenience while excluding the unwanted description field.
内容的提问来源于stack exchange,提问作者susim

