Matplotlib散点图X轴刻度标签问题及Pandas条件绘图优化咨询
Hey there! Let's tackle your two questions about plotting poker hand data—first fixing that wonky X-axis tick issue, then streamlining your conditional plotting workflow.
一、X轴刻度标签异常的修复方案
The root cause of your tick problem is likely a mismatch between your manual tick settings and the actual range of your RFI data. You set plt.xticks(x) with x = np.arange(0,1,0.1), but if your RFI values are percentages (like 0-100 instead of 0-1 decimals), this will throw off the scale entirely.
Step 1: Verify your data range first
Start by checking what your RFI values actually look like to confirm the right scale:
print(preflop_copy['RFI'].describe())
Step 2: Adjust ticks to match your data
- If
RFIis a percentage (0-100):# Set ticks from 0 to 100 in 10-point increments plt.xticks(np.arange(0, 101, 10)) - If
RFIis a decimal (0-1):# Include the 1.0 endpoint to avoid cutting off the upper range plt.xticks(np.arange(0, 1.01, 0.1)) - For zero hassle, skip manual ticks entirely: Let Matplotlib auto-generate ticks that fit your data by removing the
plt.xticks(x)line.
Bonus: Fix overlapping labels
If your tick labels are crammed together, add these lines to clean up the layout:
plt.xticks(rotation=45) plt.tight_layout() # Automatically adjusts padding to prevent label cutoff
二、无需创建新列表的Pandas条件绘图方法
Absolutely no need to make a separate filtered DataFrame! Pandas lets you combine filtering and plotting in one step using boolean indexing or the query() method for cleaner code.
Method 1: Boolean indexing (direct and straightforward)
plt.figure(figsize=(8,8)) # Filter and plot in a single line preflop_copy[preflop_copy['Hands'] > 50].plot.scatter(x="RFI", y="BB_100", ax=plt.gca()) # Add your labels and optimize layout plt.xlabel("RFI", fontsize=16) plt.ylabel("BB/100", fontsize=16) plt.tight_layout() plt.show()
Method 2: query() method (more readable for complex conditions)
Use SQL-like syntax for filtering—great if you ever need to add multiple conditions later:
plt.figure(figsize=(8,8)) # Filter with query() and plot preflop_copy.query("Hands > 50").plot.scatter(x="RFI", y="BB_100", ax=plt.gca()) plt.xlabel("RFI", fontsize=16) plt.ylabel("BB/100", fontsize=16) plt.tight_layout() plt.show()
Why this works better:
- Saves memory by avoiding intermediate DataFrames (especially useful with large datasets)
- Keeps your code logic cohesive—from data filter to plot in one flow
Full optimized code putting it all together
import pandas as pd import numpy as np import matplotlib.pyplot as plt # Data loading and preprocessing preflop = pd.read_csv("all_player_preflop_report_tourney.csv", thousands=',') preflop['Hands'] = preflop['Hands'].astype(int) # Clean up column names with a mapping (easier to read than long chained renames) column_renames = { 'BB/100':'BB_100', 'Raise First':'RFI', 'WTSD %': 'WTSD', 'All-In Adj BB/100':'adj_BB_100', 'Avg PF All-In Equity':'pf_all_in', 'CC 2Bet PF':'cc_2bet', '3Bet PF':'3bet', '2Bet PF & Call 3Bet':'2Bet_call_3Bet', 'Raise & 4Bet+ PF':'rfi_and_4bet+', '2Bet PF & Fold':'2bet_and_fold', '5Bet+ PF':'5bet+', '3Bet PF & Fold':'3bet_and_fold', 'Call Any PFR':'call_any_pfr', 'Call Steal':'call_steal', 'Call vs BTN Open':'call_btn_open', 'CC 3Bet+ PF':'cc_3bet+', 'Limp Behind':'limp_behind', 'Raise Limpers':'raise_limpers' } preflop = preflop.rename(columns=column_renames).set_index('Player') preflop_copy = preflop.copy() # Conditional plotting without intermediate DataFrames plt.figure(figsize=(8,8)) preflop_copy.query("Hands > 50").plot.scatter(x="RFI", y="BB_100", ax=plt.gca()) # Final chart polish plt.xlabel("RFI", fontsize=16) plt.ylabel("BB/100", fontsize=16) # Uncomment below if you want manual ticks (adjust range to match your RFI data) # plt.xticks(np.arange(0, 101, 10)) plt.tight_layout() plt.show()
内容的提问来源于stack exchange,提问作者jonstoeber

