如何使用Pandas替换numpy.ndarray中的缩写并保持原格式(含引号格式)
Got it, let's work through your problem step by step. First, let's unpack what's going wrong with your current approach and fix it to keep your output format identical to the input.
What's Off in Your Current Code
- Lost Original Case: By converting all text to lowercase before replacement, you're stripping out the original capitalization (like "And", "His", "I'm" turning to lowercase), which breaks the input format.
- Unnecessary DataFrame Detour: Converting to a DataFrame and back to a numpy array adds extra steps, and the quote change you're seeing is just numpy's display behavior—not an actual change to your string content.
Solution: Process the ndarray Directly (Keep Case & Format)
We can handle the abbreviation replacement directly on the numpy array, preserving original capitalization and array structure without switching to a DataFrame. Here's how:
import numpy as np import re # Your existing abbreviation dictionary abbreviations_master = { "i'm": "i am", "it's": "it is", "that's": "that is", "don't": "do not", "i'll": "i will", "i've": "i have", "we're": "we are", "didn't": "did not", "ma'am": "madam", "you're": "you are", "there's": "there is ", "let's": "let us", "they're": "they are", "can't": "can not", "he's": "he is", "doesn't": "does not", "she's": "she is", "what's": "what is", "i'd": "I would ", "haven't": "have not", "wasn't": "was not", "we'll": "we will", "won't": "will not", "it'll": "it will", "we've": "we have", "wouldn't": "would not", "that'd": "that would ", "you've": "you have", "couldn't": "could not", "that'll": "that will", "y'all": "you all", "isn't": "is not", "it'd": "it would", "would've": "would have", "'cause": "because", "hasn't": "has not", "they've": "they have", "you'll": "you will", "here's": "here is", "name's": "name is", "shouldn't": "should not", "wife's": "?", "driver's": "?", "they'll": "they will", "everything's": "?", "husband's": "?", "there'll": "there will", "should've": "should have", "we'd": "we would", "'bout": "about", "she'll": "she will", "he'll": "he will", "you'd": "you would", "one's": "?", "who's": "who has", "weren't": "were not", "aren't": "are not", "how's": "how is", "how're": "how are", "hadn't": "had not" } # Build a regex pattern to match abbreviations (word boundaries + ignore case) pattern = re.compile( r'\b(' + '|'.join(re.escape(key) for key in abbreviations_master.keys()) + r')\b', flags=re.IGNORECASE ) # Define a replacement function that preserves original capitalization def replace_match(match): original_abbrev = match.group(0) lower_key = original_abbrev.lower() replacement = abbreviations_master[lower_key] # Match the capitalization of the original abbreviation if original_abbrev[0].isupper(): return replacement.capitalize() return replacement.lower() # Your original numpy array X_trying = np.array([ [" And my account number His Okay It is Arrow My name with a K Last name Is and another phone numbers That's okay it's just number Yes <unk> at Gmail Dot com that is a lower "], ["Hi Amber I'm relocating so I need a insurance card for my car First name <unk> last name is D key No for brand new isn't"] ], dtype='<U97064') # Apply the replacement to every element in the array X_processed = np.vectorize(lambda text: pattern.sub(replace_match, text))(X_trying) # Check the result print(X_processed)
Key Notes:
- Preserved Format: This code keeps the original capitalization (e.g., "I'm" becomes "I am", "That's" becomes "That is", while lowercase "it's" stays "it is") and maintains the exact array shape/dtype of your input.
- Quote "Issue" Explained: The switch between single/double quotes in numpy's output is just a display choice—numpy uses whichever quote won't require escaping in the string. The actual string content does NOT include these quotes, so your downstream code won't be affected. If you need to verify, just access a string directly (e.g.,
X_processed[0][0]) and you'll see the raw text without any wrapping quotes.
If you really need all-lowercase text (like your original approach), you can adjust the replacement function to return all lowercase, but the core logic of processing the ndarray directly still holds.
内容的提问来源于stack exchange,提问作者user2543622
相关产品推荐
相关产品推荐

