Pandas替换列值并基于公式绘图,CSV按z分组绘制x/y与z的关系
Got it, let's break down these two Pandas data processing and visualization tasks with clear, actionable steps and code examples. I'll cover each requirement in detail below.
1. Replace Old Column Values & Visualize Derived New Columns
First up: updating existing column values in a DataFrame and creating visualizations for the new columns you compute using your specified equations.
Step-by-Step Walkthrough
- Import core libraries: You'll need
pandasfor data manipulation andmatplotlib.pyplot(orseaborn) for plotting. - Load your dataset: Use
pd.read_csv()(or the right method for your data source) to pull in your data. - Replace old values: Use
replace()for simple value mappings, orapply()with a custom function if you need more complex logic. - Calculate your new column: Plug your equation into a new column using Pandas vectorized operations (way faster than looping!).
- Visualize the results: Pick a plot type that highlights the relationship your new column shows—bar, line, scatter, etc.
Example Code
import pandas as pd import matplotlib.pyplot as plt # Load your actual data (swap this path with your file) df = pd.read_csv('your_dataset.csv') # Example: Replace values in a 'status' column value_map = {'inactive': 'archived', 'active': 'current'} df['status'] = df['status'].replace(value_map) # Example: Compute new column using an equation (adjust to your actual formula) df['calculated_col'] = (df['col_a'] * 1.5) + (df['col_b'] ** 2) # Plot: Average calculated value by updated status plt.figure(figsize=(10, 6)) df.groupby('status')['calculated_col'].mean().plot(kind='bar', color='teal') plt.title('Average Calculated Value by Status') plt.xlabel('Status') plt.ylabel('Average Calculated Value') plt.xticks(rotation=0) plt.tight_layout() plt.show()
2. Group by Column z & Plot x/y Relationships
For your CSV with z, x, y columns (where decimals use commas), we'll first fix the decimal parsing, then group by z to visualize how x and y interact across different z values.
Step-by-Step Walkthrough
- Load data correctly: Use
decimal=','inpd.read_csv()to convert comma-separated decimals (like1,75181E-07) into valid floats. - Group by
z: Split the data into subsets based on each uniquezvalue. - Plot for each group: Choose between subplots (to examine each
zindividually) or an overlay plot (to compare trends acrosszvalues).
Example Code
import pandas as pd import matplotlib.pyplot as plt # Load CSV, handling comma decimals df = pd.read_csv('your_zxy_data.csv', decimal=',') # Option 1: Subplots for each unique z value unique_z_values = df['z'].unique() num_plots = len(unique_z_values) fig, axes = plt.subplots(nrows=(num_plots + 1) // 2, ncols=2, figsize=(14, 8)) axes = axes.flatten() for idx, z_val in enumerate(unique_z_values): group_data = df[df['z'] == z_val] axes[idx].scatter(group_data['x'], group_data['y'], color='coral', alpha=0.6) axes[idx].plot(group_data['x'], group_data['y'], color='navy', linewidth=1.5) axes[idx].set_title(f'Group: z = {z_val}') axes[idx].set_xlabel('x') axes[idx].set_ylabel('y') # Hide empty subplots if needed for ax in axes[num_plots:]: ax.axis('off') plt.tight_layout() plt.show() # Option 2: Overlay all z groups in one plot for comparison plt.figure(figsize=(10, 6)) for z_val, group in df.groupby('z'): plt.plot(group['x'], group['y'], label=f'z = {z_val}', linewidth=2) plt.xlabel('x') plt.ylabel('y') plt.title('x vs y Across Different z Groups') plt.legend() plt.tight_layout() plt.show()
Pro Tip
Don't skip the decimal=',' parameter! Without it, Pandas will read your x and y values as strings instead of numbers, which breaks all your plotting and calculations.
内容的提问来源于stack exchange,提问作者Kumar

