如何用逗号分隔指定格式数据?TensorFlow练习Pandas加载数据集遇问题
Hey there, let's tackle your two questions step by step:
Your data uses semicolons (;) as the delimiter, with double quotes wrapping some fields. To switch it to comma-separated format, you have two reliable options:
Option 1: Use Pandas to read and re-save
First load the semicolon-separated data into a DataFrame, then export it with commas as the delimiter:import pandas as pd # Read the original semicolon-separated file df = pd.read_csv('your_data_file.csv', sep=';', quotechar='"') # Save as a new comma-separated CSV df.to_csv('comma_separated_data.csv', sep=',', index=False)Option 2: Raw string replacement (for text content)
If you're working directly with the raw text (not a file), you can safely replace semicolons with commas—since your fields are quoted and don't contain semicolons inside, this won't break your data structure:raw_data = '(fixed acidity;"volatile acidity";"citric acid";...)' # Remove surrounding parentheses first if needed cleaned_data = raw_data.strip('()') comma_separated = cleaned_data.replace(';', ',') print(comma_separated) # Output: fixed acidity,"volatile acidity","citric acid",...
You’re exactly right that pd.read_csv is the go-to tool here. The critical details are specifying the correct delimiter and forcing the quality column to be a float (since your regression model expects floating-point values). Here’s a complete, actionable example:
import pandas as pd # Replace 'wine_data.csv' with your actual file path df = pd.read_csv( 'wine_data.csv', sep=';', # The wine dataset uses semicolons instead of commas dtype={'quality': float} # Convert 'quality' from integer to float ) # Verify the data types to confirm it worked print(df.dtypes)
Quick breakdown of the key parameters:
sep=';': Tells Pandas to split columns on semicolons, which matches your dataset’s actual format.dtype={'quality': float}: Explicitly sets thequalitycolumn’s data type to float. Even if the original values are integers, this conversion ensures compatibility with your regression model.
Once you have the DataFrame, you can easily prepare it for TensorFlow like this:
import tensorflow as tf # Extract features (all columns except 'quality') X = df.drop('quality', axis=1).values # Extract target variable (already formatted as float) y = df['quality'].values # Build your regression model model = tf.keras.Sequential([ tf.keras.layers.Dense(64, activation='relu', input_shape=(X.shape[1],)), tf.keras.layers.Dense(1) # Single output node for regression ])
内容的提问来源于stack exchange,提问作者Daniel Rangel Martinez

