如何在DataFrame的concatenar列正确应用正则表达式提取内容?
Hey there! Let's sort out this regex issue you're facing. The error you're seeing (TypeError: expected string or bytes-like object) happens because Python's built-in re.findall() function only works with individual strings, but you're passing an entire Pandas Series (the concatenar column) directly to it. Pandas has its own string handling methods that work seamlessly with Series, so let's use those instead.
Why Your Current Code Fails
df['concatenar'] is a Series object, not a single string. re.findall() doesn't know how to process this collection—it expects one string at a time. That's why you get the type error.
Solution 1: Use Pandas str.replace() to Remove Unwanted Date Parts
Looking at your desired output, you want to keep the number + word portion and strip off the trailing date string. The date follows the pattern digits-digits-digits, so we can target that directly with str.replace():
import pandas as pd # Your original DataFrame df = pd.DataFrame({ 'concatenar': ['1 BBVA2018-03-2020', '5 santander2018-03-2020'], 'buy_sell': ['sell', 'buy'] }) # Remove the date segment from each string df['concatenar'] = df['concatenar'].str.replace(r'\d+-\d+-\d+', '')
This will modify the concatenar column exactly as you want:
| concatenar | buy_sell |
|---|---|
| 1 BBVA | sell |
| 5 santander | buy |
Solution 2: Use str.extract() to Pull Exact Matches
If you want to strictly extract the parts that match your regex (adjusted to account for both uppercase-only words like BBVA and lowercase words like santander), use str.extract() with a grouped regex:
# Extract number + space + letter sequence (matches BBVA and santander) df['concatenar'] = df['concatenar'].str.extract(r'(\d+\s+[A-Za-z]+)')
The regex (\d+\s+[A-Za-z]+) breaks down to:
\d+: Match one or more digits\s+: Match one or more spaces[A-Za-z]+: Match one or more letters (uppercase or lowercase)- The parentheses capture the entire group as the value we want to keep.
Solution 3: Using re.findall() with apply() (If You Prefer)
If you really want to use Python's re module, you can use apply() to run re.findall() on each individual string in the Series. Just make sure to handle cases where matches might be missing:
import re def extract_target_text(s): # Get the number part num = re.findall(r'\d+', s)[0] # Get the letter part (adjust regex to match your needs) text = re.findall(r'[A-Za-z]+', s)[0] return f"{num} {text}" # Apply the function to each element in the column df['concatenar'] = df['concatenar'].apply(extract_target_text)
Note: This method is less efficient for large DataFrames compared to Pandas' native str methods, since it processes each row individually.
内容的提问来源于stack exchange,提问作者JamesHudson81

