You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas替换numpy.ndarray中的缩写并保持原格式(含引号格式)

Got it, let's work through your problem step by step. First, let's unpack what's going wrong with your current approach and fix it to keep your output format identical to the input.

What's Off in Your Current Code

  1. Lost Original Case: By converting all text to lowercase before replacement, you're stripping out the original capitalization (like "And", "His", "I'm" turning to lowercase), which breaks the input format.
  2. Unnecessary DataFrame Detour: Converting to a DataFrame and back to a numpy array adds extra steps, and the quote change you're seeing is just numpy's display behavior—not an actual change to your string content.

Solution: Process the ndarray Directly (Keep Case & Format)

We can handle the abbreviation replacement directly on the numpy array, preserving original capitalization and array structure without switching to a DataFrame. Here's how:

import numpy as np
import re

# Your existing abbreviation dictionary
abbreviations_master = {
    "i'm": "i am",
    "it's": "it is",
    "that's": "that is",
    "don't": "do not",
    "i'll": "i will",
    "i've": "i have",
    "we're": "we are",
    "didn't": "did not",
    "ma'am": "madam",
    "you're": "you are",
    "there's": "there is ",
    "let's": "let us",
    "they're": "they are",
    "can't": "can not",
    "he's": "he is",
    "doesn't": "does not",
    "she's": "she is",
    "what's": "what is",
    "i'd": "I would ",
    "haven't": "have not",
    "wasn't": "was not",
    "we'll": "we will",
    "won't": "will not",
    "it'll": "it will",
    "we've": "we have",
    "wouldn't": "would not",
    "that'd": "that would ",
    "you've": "you have",
    "couldn't": "could not",
    "that'll": "that will",
    "y'all": "you all",
    "isn't": "is not",
    "it'd": "it would",
    "would've": "would have",
    "'cause": "because",
    "hasn't": "has not",
    "they've": "they have",
    "you'll": "you will",
    "here's": "here is",
    "name's": "name is",
    "shouldn't": "should not",
    "wife's": "?",
    "driver's": "?",
    "they'll": "they will",
    "everything's": "?",
    "husband's": "?",
    "there'll": "there will",
    "should've": "should have",
    "we'd": "we would",
    "'bout": "about",
    "she'll": "she will",
    "he'll": "he will",
    "you'd": "you would",
    "one's": "?",
    "who's": "who has",
    "weren't": "were not",
    "aren't": "are not",
    "how's": "how is",
    "how're": "how are",
    "hadn't": "had not"
}

# Build a regex pattern to match abbreviations (word boundaries + ignore case)
pattern = re.compile(
    r'\b(' + '|'.join(re.escape(key) for key in abbreviations_master.keys()) + r')\b',
    flags=re.IGNORECASE
)

# Define a replacement function that preserves original capitalization
def replace_match(match):
    original_abbrev = match.group(0)
    lower_key = original_abbrev.lower()
    replacement = abbreviations_master[lower_key]
    
    # Match the capitalization of the original abbreviation
    if original_abbrev[0].isupper():
        return replacement.capitalize()
    return replacement.lower()

# Your original numpy array
X_trying = np.array([
    [" And my account number His Okay It is Arrow My name with a K Last name Is and another phone numbers That's okay it's just number Yes <unk> at Gmail Dot com that is a lower "],
    ["Hi Amber I'm relocating so I need a insurance card for my car First name <unk> last name is D key No for brand new isn't"]
], dtype='<U97064')

# Apply the replacement to every element in the array
X_processed = np.vectorize(lambda text: pattern.sub(replace_match, text))(X_trying)

# Check the result
print(X_processed)

Key Notes:

  1. Preserved Format: This code keeps the original capitalization (e.g., "I'm" becomes "I am", "That's" becomes "That is", while lowercase "it's" stays "it is") and maintains the exact array shape/dtype of your input.
  2. Quote "Issue" Explained: The switch between single/double quotes in numpy's output is just a display choice—numpy uses whichever quote won't require escaping in the string. The actual string content does NOT include these quotes, so your downstream code won't be affected. If you need to verify, just access a string directly (e.g., X_processed[0][0]) and you'll see the raw text without any wrapping quotes.

If you really need all-lowercase text (like your original approach), you can adjust the replacement function to return all lowercase, but the core logic of processing the ndarray directly still holds.

内容的提问来源于stack exchange,提问作者user2543622

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 01:07:37