使用pd.read_fwf时如何为重复的Filler列名生成后缀?
pd.read_fwf Got it, let's tackle that duplicate column name error you're hitting with pd.read_fwf! The key here is to generate unique names for those repeated "Filler" columns before passing them to the function.
Step 1: Generate Unique Column Names
We can use pandas' grouping and counting tools to add numeric suffixes to duplicate "Filler" entries in your col_widths DataFrame. Here's a clean, straightforward way to do it:
import pandas as pd # Your existing column width DataFrame col_widths = pd.DataFrame({ 'name': ['Filler', 'Col A', 'Filler', 'Col B'], 'width': [2, 8, 4, 6] }) # Create unique names by adding numeric suffixes to duplicate "Filler" columns col_widths['unique_name'] = col_widths['name'].where( col_widths['name'] != 'Filler', col_widths['name'] + '_' + (col_widths.groupby('name').cumcount() + 1).astype(str) )
Let me break this down:
groupby('name').cumcount()counts occurrences of each column name starting from 0. Adding 1 shifts it to start at 1, so we get_1,_2, etc.- We only apply this suffix logic to "Filler" columns, leaving your actual data columns (like "Col A", "Col B") with their original names.
Step 2: Read the Fixed-Width Data
Now pass the unique names and widths to pd.read_fwf:
# Extract the unique names and widths as lists to feed into read_fwf unique_names = col_widths['unique_name'].tolist() column_widths = col_widths['width'].tolist() # Read your fixed-width data file with the unique column names df = pd.read_fwf('your_data_file.txt', widths=column_widths, names=unique_names)
Optional: Remove Filler Columns (If Needed)
If you don't want those Filler columns cluttering your final DataFrame, you can drop them in one line:
df = df.drop(columns=[col for col in df.columns if 'Filler' in col])
That's it! This will eliminate the ValueError from duplicate names and let you read your fixed-width data smoothly.
内容的提问来源于stack exchange,提问作者cpage

