如何基于27000行CSV文件创建指定样式的Jitter Plot?
Got it, let's build that jitter plot you need for your 27,000-row CSV of 2017 Brazilian senator expenses. We'll use Python with Pandas (for data handling) and Seaborn/Matplotlib (for plotting)—this stack handles large datasets smoothly and lets you replicate the clean, informative style you're aiming for.
Step 1: Install Required Libraries
First, make sure you have these packages installed (run this in your terminal):
pip install pandas seaborn matplotlib
Step 2: Load and Clean Your Data
We'll start by reading the CSV, filtering for 2017 entries, ensuring the refund amount is numeric, and sorting months in chronological order to keep the plot intuitive.
import pandas as pd import seaborn as sns import matplotlib.pyplot as plt from matplotlib.ticker import FuncFormatter # Load your CSV file (replace with your actual file path) df = pd.read_csv("your_expense_data.csv") # Filter for 2017 data (adjust if your year column has a different name, like "YEAR") df_2017 = df[df["ANO"] == 2017] # Convert refund amount to numeric (in case it's stored as text with currency symbols) df_2017["VALOR_REEMBOLSADO"] = pd.to_numeric(df_2017["VALOR_REEMBOLSADO"], errors="coerce") # Define chronological order of Brazilian months (adjust if your MES column uses numbers 1-12 instead) month_order = ["Janeiro", "Fevereiro", "Março", "Abril", "Maio", "Junho", "Julho", "Agosto", "Setembro", "Outubro", "Novembro", "Dezembro"] df_2017["MES"] = pd.Categorical(df_2017["MES"], categories=month_order, ordered=True)
Step 3: Build the Jitter Plot
We'll use seaborn.stripplot (ideal for large datasets) with jitter to spread overlapping points. Transparency helps handle the high volume of data, and we'll format the Y-axis as Brazilian Reais for clarity.
# Set a clean, professional plot style sns.set_style("whitegrid") # Create the plot with adjusted size for readability plt.figure(figsize=(12, 6)) jitter_plot = sns.stripplot( x="MES", y="VALOR_REEMBOLSADO", data=df_2017, jitter=0.3, # Controls how much points spread horizontally alpha=0.2, # Transparency prevents overcrowding size=3, # Smaller point size works better for 27k rows color="#2c3e50" ) # Format Y-axis to show Brazilian Reais def format_reais(value, tick_number): return f"R$ {value:,.0f}" jitter_plot.yaxis.set_major_formatter(FuncFormatter(format_reais)) # Customize plot labels and title plt.title("2017 Brazilian Senator Refund Expenses by Month", fontsize=14, pad=20) plt.xlabel("Month", fontsize=12) plt.ylabel("Refund Amount (R$)", fontsize=12) # Rotate X-axis labels so they don't overlap plt.xticks(rotation=45) # Remove unnecessary spines for a cleaner look sns.despine() # Adjust layout to prevent label cutoff plt.tight_layout() # Display the plot plt.show()
Quick Adjustments for Your Dataset:
- If your
MEScolumn uses numeric values (1-12), update themonth_orderlist to[1, 2, 3, ..., 12]. - If extreme outliers skew the Y-axis, add
ylim=(0, your_upper_limit)toplt.figure()to focus on the main distribution (just note this hides some data points). - For slightly smoother rendering, you can swap
stripplotwithswarmplot, butstripplotis faster for large datasets like yours.
内容的提问来源于stack exchange,提问作者agccaesar

