如何将NLTK地址预处理与向量化代码改写为可复用Python函数
Got it, let's turn your address preprocessing and vectorization code into a reusable, flexible function that you can call whenever you need to process address data. I've fixed a small typo in your original code (""join → "".join) and made the function configurable so you can adjust parameters like custom stopwords, feature count, and n-gram range easily.
import nltk import string import pandas as pd from sklearn.feature_extraction.text import CountVectorizer def preprocess_and_vectorize_addresses(data, addr_col='Adj_Addr', custom_stop_words=['st','rd','kwun tong','kwai chung','kwun','tong'], max_features=200, ngram_range=(1,3)): # Load default English stopwords and merge with custom ones stopwords = nltk.corpus.stopwords.words('english') combined_stop_words = stopwords + custom_stop_words # Define a helper function to clean a single address string def clean_address(addr): # Convert to lowercase and split into tokens tokens = addr.lower().split() # Remove digits tokens_no_digits = [token for token in tokens if not token.isdigit()] # Remove punctuation (strip punctuation from each token) tokens_no_punct = [token.strip(string.punctuation) for token in tokens_no_digits] # Filter out empty strings and stopwords cleaned_tokens = [token for token in tokens_no_punct if token and token not in combined_stop_words] # Join back into a single string return ' '.join(cleaned_tokens) # Apply cleaning to the address column data['Clean_addr'] = data[addr_col].apply(clean_address) # Initialize CountVectorizer and fit-transform cleaned addresses cv = CountVectorizer(max_features=max_features, analyzer='word', ngram_range=ngram_range) cv_addr = cv.fit_transform(data['Clean_addr']) # Convert sparse matrix to DataFrame columns and add to original data for i, col in enumerate(cv.get_feature_names_out()): # Uses updated sklearn method (>=1.0) data[col] = pd.SparseSeries(cv_addr[:, i].toarray().ravel(), fill_value=0) # Optional: Uncomment to drop the temporary Clean_addr column # data = data.drop('Clean_addr', axis=1) return data
Key improvements & notes:
- Reusable: Call this function with different DataFrames or tweak parameters without rewriting core logic.
- Readable helper: The
clean_addresshelper breaks down preprocessing into clear steps, making it easy to modify if you need to add/remove cleaning rules. - Fixed bug: Corrected the invalid
""joinsyntax from your original code that would throw an error. - Configurable: Adjust custom stopwords, feature limits, or n-gram ranges directly when calling the function.
- Sklearn-compatible: Uses
get_feature_names_out()instead of the deprecatedget_feature_names()for newer sklearn versions.
How to use it:
# Basic call with default settings processed_data = preprocess_and_vectorize_addresses(your_input_dataframe) # Customized call example custom_stops = ['ave', 'blvd', 'hk'] processed_data = preprocess_and_vectorize_addresses(your_input_dataframe, addr_col='Raw_Address', custom_stop_words=custom_stops, max_features=300, ngram_range=(1,2))
Quick setup note:
If you haven't already, download the NLTK stopwords first to avoid errors:
nltk.download('stopwords')
内容的提问来源于stack exchange,提问作者Snehal pankaj
相关产品推荐
相关产品推荐

