You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析非规范amenities列数据并转字符串列表的问题

Got it, let's walk through how to clean that messy amenities column and get it ready for one-hot encoding—this is a super common pain point in feature engineering when dealing with unstructured or inconsistently formatted data.

Step 1: Clean the Amenities Column

First, we need to turn those curly-brace-wrapped, mixed-quote entries into clean lists of amenities. Here's a step-by-step function to handle this:

import pandas as pd
from sklearn.preprocessing import MultiLabelBinarizer

def clean_amenities(raw_value):
    # Convert to string first to handle any non-string values (like NaN)
    amenity_str = str(raw_value)
    
    # Strip off the curly braces from start and end
    stripped = amenity_str.strip('{}')
    
    # Remove all double quotes (since some entries have them, some don't)
    no_quotes = stripped.replace('"', '')
    
    # Split into individual items, strip whitespace, and filter out empty strings
    # Handles messy cases like ", , Kitchen" that might pop up from bad formatting
    cleaned_list = [item.strip() for item in no_quotes.split(',') if item.strip()]
    
    return cleaned_list

Apply this function to your DataFrame's amenities column:

# Replace 'your_dataframe' with your actual DataFrame name
df['cleaned_amenities'] = df['amenities'].apply(clean_amenities)

What this does:

  • Converts every value to a string to avoid errors from NaN or non-string entries
  • Removes the outer {} so we can access the actual amenity values
  • Normalizes all entries by stripping quotes (since some are quoted and others aren't)
  • Splits on commas, cleans up whitespace, and drops empty entries to eliminate junk values
Step 2: One-Hot Encode the Cleaned Lists

Now that we have proper lists of amenities, we can use MultiLabelBinarizer from scikit-learn to create one-hot encoded features:

# Initialize the binarizer
mlb = MultiLabelBinarizer()

# Fit it to our cleaned lists and transform the data
amenities_one_hot = mlb.fit_transform(df['cleaned_amenities'])

# Convert the result to a DataFrame with meaningful column names
one_hot_df = pd.DataFrame(
    amenities_one_hot,
    columns=mlb.classes_,  # Uses the amenity names as column headers
    index=df.index         # Keep the same index as original data
)

# Merge the one-hot features back into your original DataFrame
df = pd.concat([df, one_hot_df], axis=1)

Quick notes:

  • If you have amenities with special characters (like Free Wi-Fi), the column names will match exactly
  • Empty lists (from entries like {} or NaN) will result in all 0s for the one-hot columns, which is correct for missing data
  • You can drop the original amenities and intermediate cleaned_amenities columns if you don't need them anymore

内容的提问来源于stack exchange,提问作者Kylo Ren

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:48:56