Matplotlib报错ValueError: color参数需为每个数据集指定一种颜色
ValueError: color kwarg must have one color per dataset in Seaborn/Matplotlib Hey there, let's sort out this error you're running into! That specific message pops up because Seaborn thinks you're trying to plot multiple datasets at once, but you only provided a single color. The most common cause here is that your data selection is returning a DataFrame instead of a single-column Series—and when distplot gets a DataFrame, it treats each column as a separate dataset.
Here are two solid solutions to fix this:
Solution 1: Fix your data indexing to get a single Series
Your current chained indexing (featureSet[featureSet['label']=='0']['len of url']) can sometimes return a DataFrame instead of a Series, especially with column names that have spaces. Switch to using loc to explicitly grab the single column you need:
import matplotlib.pyplot as plt import pandas as pd import seaborn as sns import pickle as pkl from __future__ import division sns.set(style="darkgrid") # Use .loc to safely get a single-column Series sns.distplot(featureSet.loc[featureSet['label']=='0', 'len of url'], color='green', label='Benign URLs') sns.distplot(featureSet.loc[featureSet['label']=='1', 'len of url'], color='red', label='Phishing URLs') plt.title('Url Length Distribution') plt.legend(loc='upper right') plt.xlabel('Length of URL') plt.show()
Solution 2: Use Seaborn's newer, recommended API (preferred!)
distplot has been deprecated since Seaborn version 0.11.0, so switching to histplot (with kde=True to keep the density curve like distplot had) is a better long-term fix—it's more explicit and avoids old API quirks:
import matplotlib.pyplot as plt import pandas as pd import seaborn as sns import pickle as pkl from __future__ import division sns.set(style="darkgrid") # Replace distplot with histplot + kde=True for the same visual sns.histplot(featureSet.loc[featureSet['label']=='0', 'len of url'], color='green', label='Benign URLs', kde=True) sns.histplot(featureSet.loc[featureSet['label']=='1', 'len of url'], color='red', label='Phishing URLs', kde=True) plt.title('Url Length Distribution') plt.legend(loc='upper right') plt.xlabel('Length of URL') plt.show()
One quick check to add: make sure your len of url column doesn't have missing values. If it does, clean it up first with:
featureSet = featureSet.dropna(subset=['len of url'])
内容的提问来源于stack exchange,提问作者Willze Fortner

