You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Seaborn散点图Y轴刻度点重叠问题及样式优化需求

Customizing Seaborn Scatter Plot for Your Yearly Participant Data

Hey there! Let's work through styling your Seaborn scatter plot to match your target example. Since you didn’t share the exact example, I’ll cover the most common requested customizations for this kind of yearly participant data—you can tweak these to fit your specific needs.

First, let’s get your data into a usable pandas DataFrame (I’ve formatted your raw data here):

import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt

# Load your data
data = {
    "Year": [2019,2019,2018,2018,2018,2018,2018,2018,2018,2018,2018,2018,
             2017,2017,2017,2017,2017,2017,2016,2016,2016,2016,2016,2016,
             2015,2015,2015,2015,2015,2015,2015,2015],
    "UserStudy_NumParticipants": [40,30,10,2,623,10,43,44,36,60,22,129,
                                 20,16,18,20,30,31,40,250,28,40,39,171,
                                 40,2,150,30,45,225,50,5]
}
df = pd.DataFrame(data)

Core Styling Customizations

Here are step-by-step tweaks to get a polished plot:

1. Set a Clean Base Theme

Start with a Seaborn theme that aligns with common professional styles:

sns.set_theme(style="whitegrid", font_scale=1.2)  # Adjust font scale for readability

2. Scatter Plot with Jitter & Distinction

Avoid overlapping points and make year groups clear:

fig, ax = plt.subplots(figsize=(10, 6))

# Scatter plot with jitter (prevents overlapping points in the same year)
sns.scatterplot(
    data=df,
    x="Year",
    y="UserStudy_NumParticipants",
    hue="Year",  # Color-code points by year
    s=100,  # Adjust point size
    alpha=0.7,  # Add transparency to see overlapping points
    edgecolor="black",  # Add sharp edges to points for clarity
    ax=ax
)

3. Add Distribution Context (Optional)

If your example includes elements like box plots to show yearly distributions, layer them behind the scatter:

# Add semi-transparent box plots to show spread/hidden outliers
sns.boxplot(
    data=df,
    x="Year",
    y="UserStudy_NumParticipants",
    boxprops={"alpha": 0.3},  # Keep scatter points visible
    showfliers=False,  # Hide outliers here since scatter shows them
    ax=ax
)

4. Polish Axis & Labels

Tweak titles, labels, and grids for clarity:

# Customize text elements
ax.set_title("User Study Participant Counts by Year", pad=20, fontweight="bold")
ax.set_xlabel("Year", labelpad=15)
ax.set_ylabel("Number of Participants", labelpad=15)

# Remove redundant legend (hue matches x-axis categories)
ax.legend().remove()

# Refine grid lines
ax.grid(True, axis="y", linestyle="--", alpha=0.7)

# Adjust layout to prevent label cutoff
plt.tight_layout()

plt.show()

Alternative: Add Mean Trend Line

If your target example includes a trend line for yearly averages, replace the box plot with this:

# Calculate yearly average participants
year_means = df.groupby("Year")["UserStudy_NumParticipants"].mean().reset_index()

# Plot mean line over scatter
sns.lineplot(
    data=year_means,
    x="Year",
    y="UserStudy_NumParticipants",
    color="crimson",
    marker="o",
    linewidth=2,
    ax=ax
)

Extra Tweaks for Specific Styles

  • Log Scale for Outliers: If your example uses a log y-axis to handle large values like 623, add ax.set_yscale("log")
  • Custom Colors: Use palette="viridis" or a custom list like palette=["#2ecc71", "#3498db", "#e74c3c", "#f39c12", "#9b59b6"] in scatterplot()
  • Point Shapes: Change marker style with marker="s" (square) or marker="^" (triangle)

If you can share more details about your target example (like specific colors, elements, or layout), I can refine this further!

内容的提问来源于stack exchange,提问作者vinnyalvs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 21:57:55