如何在Python中读取文本数据集并自定义列名以类CSV形式处理?
No problem! Your dataset uses spaces as delimiters instead of commas, but we can handle it just like a CSV file with Python—here's a straightforward, step-by-step approach using the most common tools:
Core Idea
Since your data is split by whitespace (not commas), we just need to tell our data-reading tool to split on spaces instead of commas, then attach your custom column names. The easiest way is with pandas (the standard library for tabular data), but I’ll also include a vanilla Python option if you don’t want to use external libraries.
Option 1: Using Pandas (Recommended)
Pandas’ read_csv function works perfectly for space-separated data—we just need to tweak a few parameters.
Step 1: Define Custom Columns
First, create a list of column names. Your dataset has 26 columns per row, so here’s a generic list (replace these with meaningful names if you know what each column represents):
custom_columns = [ 'col1', 'col2', 'col3', 'col4', 'col5', 'col6', 'col7', 'col8', 'col9', 'col10', 'col11', 'col12', 'col13', 'col14', 'col15', 'col16', 'col17', 'col18', 'col19', 'col20', 'col21', 'col22', 'col23', 'col24', 'col25', 'col26' ]
Step 2: Read the Data
If your data is saved in a text file (e.g., data.txt):
import pandas as pd # Read the space-separated data df = pd.read_csv( 'data.txt', sep='\s+', # Split on any number of spaces/tabs header=None, # The data has no built-in header names=custom_columns # Apply our custom column names ) # Verify the result print(df.head())
If your data is stored as a string (like the snippet you provided):
import pandas as pd from io import StringIO # Your raw data string data_str = """1 1 -0.0007 -0.0004 100.0 518.67 641.82 1589.70 1400.60 14.62 21.61 554.36 2388.06 9046.19 1.30 47.47 521.66 2388.02 8138.62 8.4195 0.03 392 2388 100.00 39.06 23.4190 1 2 0.0019 -0.0003 100.0 518.67 642.15 1591.82 1403.14 14.62 21.61 553.75 2388.04 9044.07 1.30 47.49 522.28 2388.07 8131.49 8.4318 0.03 392""" # Treat the string as a file-like object with StringIO df = pd.read_csv( StringIO(data_str), sep='\s+', header=None, names=custom_columns ) print(df)
Option 2: Vanilla Python (No Pandas)
If you prefer to stick with built-in tools, use the csv module. We just need to filter out empty strings caused by multiple spaces:
import csv from io import StringIO data_str = """1 1 -0.0007 -0.0004 100.0 518.67 641.82 1589.70 1400.60 14.62 21.61 554.36 2388.06 9046.19 1.30 47.47 521.66 2388.02 8138.62 8.4195 0.03 392 2388 100.00 39.06 23.4190 1 2 0.0019 -0.0003 100.0 518.67 642.15 1591.82 1403.14 14.62 21.61 553.75 2388.04 9044.07 1.30 47.49 522.28 2388.07 8131.49 8.4318 0.03 392""" custom_columns = [f'col_{i+1}' for i in range(26)] # Read and clean the data rows = [] with StringIO(data_str) as f: reader = csv.reader(f, delimiter=' ') for row in reader: # Remove empty strings from multiple spaces cleaned_row = [item for item in row if item] rows.append(cleaned_row) # Convert to a list of dictionaries with column names data_with_columns = [dict(zip(custom_columns, row)) for row in rows] # Print the first row to check print(data_with_columns[0])
Key Tips
sep='\s+'is essential: it matches any number of whitespace characters, so it handles cases where there are multiple spaces between values.header=Noneensures pandas doesn’t mistake the first data row for a header.- Replace the generic
col1,col2, etc., names with meaningful labels (e.g.,sensor_id,temperature) if you know what each column represents—this will make your analysis much easier later.
内容的提问来源于stack exchange,提问作者sumitpal0593

