H2O import_file方法使用col_names参数时抛出ValueError异常的原因咨询
col_names Fails but Manual Column Assignment Works for LibSVM Files in H2O This discrepancy boils down to a bug in how H2O 3.36.1.2 handles column name parsing for LibSVM-formatted files, specifically with the col_names parameter vs. manual column assignment:
1. Manual column assignment works because H2O correctly identifies all columns
When you run:
df1 = h2o.import_file(path="ijcnn1.tr") df1.columns = col_names
H2O properly parses the LibSVM file into 23 total columns: 1 label column (the first value in each line) plus 22 sparse feature columns. Your col_names list (1 class name + 22 F* names) matches this 23-column count perfectly, so the assignment succeeds.
2. The col_names parameter fails due to a parsing bug
When you pass col_names directly to import_file, H2O's internal parse_setup function makes a mistake with LibSVM files: it only counts the 22 feature columns and ignores the label column. This leads it to expect a col_names list of length 22, but you're providing 23 names (including the label), hence the ValueError: length of col_names should be equal to the number of columns: 23 vs 22.
Fixes & Workarounds
You have a couple of options to resolve this:
- Stick with your working method: Continue importing the file first, then assigning column names manually. This is the most reliable approach for your current H2O version.
- Temporary workaround for
col_names: If you prefer using the parameter directly, pass only the feature column names, then rename the label column afterward:feature_names = ['F' + str(i) for i in range(22)] df2 = h2o.import_file(path="ijcnn1.tr", col_names=feature_names) df2.columns = ['class'] + feature_names - Upgrade H2O: This bug appears to be version-specific. Upgrading to a newer release (e.g., 3.38 or later) likely fixes the incorrect column count calculation for LibSVM files when using
col_names.
内容的提问来源于stack exchange,提问作者rozyang

