You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将NLTK地址预处理与向量化代码改写为可复用Python函数

Got it, let's turn your address preprocessing and vectorization code into a reusable, flexible function that you can call whenever you need to process address data. I've fixed a small typo in your original code (""join → "".join) and made the function configurable so you can adjust parameters like custom stopwords, feature count, and n-gram range easily.

import nltk
import string
import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer

def preprocess_and_vectorize_addresses(data, addr_col='Adj_Addr', 
                                      custom_stop_words=['st','rd','kwun tong','kwai chung','kwun','tong'],
                                      max_features=200, ngram_range=(1,3)):
    # Load default English stopwords and merge with custom ones
    stopwords = nltk.corpus.stopwords.words('english')
    combined_stop_words = stopwords + custom_stop_words
    
    # Define a helper function to clean a single address string
    def clean_address(addr):
        # Convert to lowercase and split into tokens
        tokens = addr.lower().split()
        # Remove digits
        tokens_no_digits = [token for token in tokens if not token.isdigit()]
        # Remove punctuation (strip punctuation from each token)
        tokens_no_punct = [token.strip(string.punctuation) for token in tokens_no_digits]
        # Filter out empty strings and stopwords
        cleaned_tokens = [token for token in tokens_no_punct if token and token not in combined_stop_words]
        # Join back into a single string
        return ' '.join(cleaned_tokens)
    
    # Apply cleaning to the address column
    data['Clean_addr'] = data[addr_col].apply(clean_address)
    
    # Initialize CountVectorizer and fit-transform cleaned addresses
    cv = CountVectorizer(max_features=max_features, analyzer='word', ngram_range=ngram_range)
    cv_addr = cv.fit_transform(data['Clean_addr'])
    
    # Convert sparse matrix to DataFrame columns and add to original data
    for i, col in enumerate(cv.get_feature_names_out()):  # Uses updated sklearn method (>=1.0)
        data[col] = pd.SparseSeries(cv_addr[:, i].toarray().ravel(), fill_value=0)
    
    # Optional: Uncomment to drop the temporary Clean_addr column
    # data = data.drop('Clean_addr', axis=1)
    
    return data

Key improvements & notes:

  • Reusable: Call this function with different DataFrames or tweak parameters without rewriting core logic.
  • Readable helper: The clean_address helper breaks down preprocessing into clear steps, making it easy to modify if you need to add/remove cleaning rules.
  • Fixed bug: Corrected the invalid ""join syntax from your original code that would throw an error.
  • Configurable: Adjust custom stopwords, feature limits, or n-gram ranges directly when calling the function.
  • Sklearn-compatible: Uses get_feature_names_out() instead of the deprecated get_feature_names() for newer sklearn versions.

How to use it:

# Basic call with default settings
processed_data = preprocess_and_vectorize_addresses(your_input_dataframe)

# Customized call example
custom_stops = ['ave', 'blvd', 'hk']
processed_data = preprocess_and_vectorize_addresses(your_input_dataframe,
                                                    addr_col='Raw_Address',
                                                    custom_stop_words=custom_stops,
                                                    max_features=300,
                                                    ngram_range=(1,2))

Quick setup note:

If you haven't already, download the NLTK stopwords first to avoid errors:

nltk.download('stopwords')

内容的提问来源于stack exchange,提问作者Snehal pankaj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:46:45