You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中遍历DataFrame每行提取两子串间的字符串

Extract Substring Between Two Strings for Every Row in a Pandas DataFrame

Hey there! Let's adapt your existing find_between() function to work across every row in your DataFrame. I'll cover two practical approaches—one that reuses your existing logic, and a faster optimized method for larger datasets.

First, recap your original function

I assume your current find_between() looks something like this (tweak it if your implementation is different):

def find_between(s, start, end):
    try:
        # Split on the start substring, grab the second segment, then split on end and take the first part
        return s.split(start)[1].split(end)[0]
    except IndexError:
        # Return empty string or None if start/end aren't found in the string
        return ""

Approach 1: Reuse your function with apply()

This is the most straightforward way to leverage your existing code. We'll apply the function to every value in your target DataFrame column:

import pandas as pd

# Example DataFrame (replace with your actual data)
df = pd.DataFrame({
    'content': [
        "Here's [START]the text we need[END] to pull out",
        "Another entry with [START]unique content[END] inside",
        "A row without matching substrings at all"
    ]
})

# Define your find_between function as above
def find_between(s, start, end):
    try:
        return s.split(start)[1].split(end)[0]
    except IndexError:
        return ""

# Apply the function to the 'content' column and store results in a new column
df['extracted_text'] = df['content'].apply(lambda x: find_between(x, '[START]', '[END]'))

print(df)

Approach 2: Faster Vectorized Method with str.extract()

For larger DataFrames, apply() can be slow because it loops through each row individually. Pandas has built-in vectorized string operations that are way more efficient. We'll use regular expressions with str.extract():

import pandas as pd

# Same example DataFrame
df = pd.DataFrame({
    'content': [
        "Here's [START]the text we need[END] to pull out",
        "Another entry with [START]unique content[END] inside",
        "A row without matching substrings at all"
    ]
})

# Use regex lookbehind/lookahead to extract text between your start and end substrings
# Replace '[START]' and '[END]' with your actual target substrings
df['extracted_text'] = df['content'].str.extract(r'(?<=\[START\])(.*?)(?=\[END\])', expand=False)

# Fill NaN values (from rows where start/end weren't found) with empty string
df['extracted_text'] = df['extracted_text'].fillna("")

print(df)

Quick Notes:

  • If your start/end substrings contain special regex characters (like ., *, [), escape them using re.escape() to avoid errors:
    import re
    start = re.escape('[START]')
    end = re.escape('[END]')
    df['extracted_text'] = df['content'].str.extract(r'(?<=' + start + ')(.*?)(?=' + end + ')', expand=False)
    
  • The vectorized method is always preferred for big datasets—it cuts down on processing time significantly.

内容的提问来源于stack exchange,提问作者lalatei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:51:06