DataFrame列值处理:提取首个元素,正则替换失效求助
Your current regex isn’t working because it’s targeting digits after the comma (\d+), but your values have letters (like D,A,C). Let’s fix this with two straightforward, reliable approaches:
Approach 1: Split on Comma and Take the First Element
This is the most intuitive method for this task. Split each string at the comma, then grab the first item from the resulting list:
df["col name"] = df["col name"].astype(str).str.split(',').str[0]
str.split(',')turns each string into a list (e.g.,'D,A,C'becomes['D', 'A', 'C'])str[0]extracts the first element of each list (so['D', 'A', 'C']simplifies to'D')
Approach 2: Fixed Regex Replace
If you prefer using regex, adjust your pattern to match any character after the comma instead of just digits. Use r',.*' which means "comma followed by zero or more of any character":
df["col name"] = df["col name"].astype(str).str.replace(r',.*', '', regex=True)
,.*matches a comma and everything that comes after it- Replacing that match with an empty string leaves only the first element before the comma
Both methods will transform your original array from:array(['A','B','C','D','D,A,C','B,A'], dtype=object)
to your desired result:array(['A','B','C','D','D','B'], dtype=object)
Pick the approach that feels most readable to you—they’ll both get the job done perfectly!
内容的提问来源于stack exchange,提问作者tbone

