You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame的concatenar列正确应用正则表达式提取内容?

Fixing the Regex Application on Your Pandas DataFrame

Hey there! Let's sort out this regex issue you're facing. The error you're seeing (TypeError: expected string or bytes-like object) happens because Python's built-in re.findall() function only works with individual strings, but you're passing an entire Pandas Series (the concatenar column) directly to it. Pandas has its own string handling methods that work seamlessly with Series, so let's use those instead.

Why Your Current Code Fails

df['concatenar'] is a Series object, not a single string. re.findall() doesn't know how to process this collection—it expects one string at a time. That's why you get the type error.

Solution 1: Use Pandas str.replace() to Remove Unwanted Date Parts

Looking at your desired output, you want to keep the number + word portion and strip off the trailing date string. The date follows the pattern digits-digits-digits, so we can target that directly with str.replace():

import pandas as pd

# Your original DataFrame
df = pd.DataFrame({
    'concatenar': ['1 BBVA2018-03-2020', '5 santander2018-03-2020'],
    'buy_sell': ['sell', 'buy']
})

# Remove the date segment from each string
df['concatenar'] = df['concatenar'].str.replace(r'\d+-\d+-\d+', '')

This will modify the concatenar column exactly as you want:

concatenarbuy_sell
1 BBVAsell
5 santanderbuy

Solution 2: Use str.extract() to Pull Exact Matches

If you want to strictly extract the parts that match your regex (adjusted to account for both uppercase-only words like BBVA and lowercase words like santander), use str.extract() with a grouped regex:

# Extract number + space + letter sequence (matches BBVA and santander)
df['concatenar'] = df['concatenar'].str.extract(r'(\d+\s+[A-Za-z]+)')

The regex (\d+\s+[A-Za-z]+) breaks down to:

  • \d+: Match one or more digits
  • \s+: Match one or more spaces
  • [A-Za-z]+: Match one or more letters (uppercase or lowercase)
  • The parentheses capture the entire group as the value we want to keep.

Solution 3: Using re.findall() with apply() (If You Prefer)

If you really want to use Python's re module, you can use apply() to run re.findall() on each individual string in the Series. Just make sure to handle cases where matches might be missing:

import re

def extract_target_text(s):
    # Get the number part
    num = re.findall(r'\d+', s)[0]
    # Get the letter part (adjust regex to match your needs)
    text = re.findall(r'[A-Za-z]+', s)[0]
    return f"{num} {text}"

# Apply the function to each element in the column
df['concatenar'] = df['concatenar'].apply(extract_target_text)

Note: This method is less efficient for large DataFrames compared to Pandas' native str methods, since it processes each row individually.


内容的提问来源于stack exchange,提问作者JamesHudson81

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:36:34