基于Bokeh的单数据列箱线图生成问题:现有分组箱线图代码无法适配的解决建议
The issue with your current code is that it's hardcoded to rely on grouping by a categorical variable. To create a box plot for a single column (like math score) showing its overall distribution, we need to adjust the function to handle a "single group" scenario where we treat the entire dataset as one category.
Here's the modified version of your function that supports both grouped and single-column box plots:
from bokeh.models import Range1d from bokeh.plotting import figure, show def box_plot(df, vals, label=None, ylabel=None, xlabel=None, title=None): # Handle single column (no grouping) case if label is None: # Create a dummy group with the entire dataset df_gb = df.groupby(lambda x: "Overall") cats = ["Overall"] else: # Original grouping logic df_gb = df.groupby(label) cats = list(df_gb.groups.keys()) cats = [str(i) for i in cats] # Compute quartiles for each group q1 = df_gb[vals].quantile(q=0.25) q2 = df_gb[vals].quantile(q=0.5) q3 = df_gb[vals].quantile(q=0.75) # Compute interquartile region and outlier bounds iqr = q3 - q1 upper_cutoff = q3 + 1.5*iqr lower_cutoff = q1 - 1.5*iqr # Find outliers for each group def outliers(group): cat = group.name outlier_inds = (group[vals] > upper_cutoff[cat]) | (group[vals] < lower_cutoff[cat]) return group[vals][outlier_inds] out = df_gb.apply(outliers).dropna() # Prepare outlier points for plotting outx = [] outy = [] for cat in cats: if cat in out and not out[cat].empty: for value in out[cat]: outx.append(cat) outy.append(value) # Adjust whiskers to exclude outliers qmin = df_gb[vals].min() qmax = df_gb[vals].max() upper = [min([x,y]) for (x,y) in zip(qmax, upper_cutoff)] lower = [max([x,y]) for (x,y) in zip(qmin, lower_cutoff)] # Build figure p = figure(sizing_mode='stretch_width', x_range=cats, height=300, toolbar_location=None) p.xgrid.grid_line_color = None p.ygrid.grid_line_width = 2 p.yaxis.axis_label = ylabel if ylabel else vals p.xaxis.axis_label = xlabel if xlabel else (label if label else "") p.title = title if title else f"Box Plot of {vals}" p.y_range.start = 0 p.title.align = 'center' # Stems p.segment(cats, upper, cats, q3, line_width=2, line_color="black") p.segment(cats, lower, cats, q1, line_width=2, line_color="black") # Boxes - use a single color for single group, or original palette for multiple groups fill_colors = ['#a50f15'] if label is None else ['#a50f15', '#de2d26', '#fb6a4a', '#fcae91', '#fee5d9'] p.rect(cats, (q3 + q1)/2, 0.5, q3 - q1, fill_color=fill_colors, alpha=0.7, line_width=2, line_color="black") # Median line p.rect(cats, q2, 0.5, 0.01, line_color="black", line_width=2) # Whisker caps p.rect(cats, lower, 0.2, 0.01, line_color="black") p.rect(cats, upper, 0.2, 0.01, line_color="black") # Outliers p.circle(outx, outy, size=6, color="black") return p
Key Changes Explained:
- Added a conditional check for
label=None: When you don't pass a categorical label, the function creates a dummy group called "Overall" containing the entire dataset. - Adjusted the fill color palette to use a single color for the single-group case (you can change this to any color you prefer).
- Improved default labels/title for better readability when no custom text is provided.
How to Use It for Single Column:
To generate a box plot for math score, simply call the function without passing a label parameter:
# Assuming df is your loaded StudentsPerformance dataset p = box_plot(df, 'math score', ylabel='Math Score', title='Distribution of Math Scores') show(p)
Why Your Previous Attempt Failed:
Setting cats = df['math score'] was incorrect because cats is meant to be the list of category labels (like group names), not the raw data values. For a single column, you only need one category label (like "Overall") to anchor the box plot on the x-axis.
内容的提问来源于stack exchange,提问作者curiouscoder

