如何在Python中遍历DataFrame每行提取两子串间的字符串
Hey there! Let's adapt your existing find_between() function to work across every row in your DataFrame. I'll cover two practical approaches—one that reuses your existing logic, and a faster optimized method for larger datasets.
First, recap your original function
I assume your current find_between() looks something like this (tweak it if your implementation is different):
def find_between(s, start, end): try: # Split on the start substring, grab the second segment, then split on end and take the first part return s.split(start)[1].split(end)[0] except IndexError: # Return empty string or None if start/end aren't found in the string return ""
Approach 1: Reuse your function with apply()
This is the most straightforward way to leverage your existing code. We'll apply the function to every value in your target DataFrame column:
import pandas as pd # Example DataFrame (replace with your actual data) df = pd.DataFrame({ 'content': [ "Here's [START]the text we need[END] to pull out", "Another entry with [START]unique content[END] inside", "A row without matching substrings at all" ] }) # Define your find_between function as above def find_between(s, start, end): try: return s.split(start)[1].split(end)[0] except IndexError: return "" # Apply the function to the 'content' column and store results in a new column df['extracted_text'] = df['content'].apply(lambda x: find_between(x, '[START]', '[END]')) print(df)
Approach 2: Faster Vectorized Method with str.extract()
For larger DataFrames, apply() can be slow because it loops through each row individually. Pandas has built-in vectorized string operations that are way more efficient. We'll use regular expressions with str.extract():
import pandas as pd # Same example DataFrame df = pd.DataFrame({ 'content': [ "Here's [START]the text we need[END] to pull out", "Another entry with [START]unique content[END] inside", "A row without matching substrings at all" ] }) # Use regex lookbehind/lookahead to extract text between your start and end substrings # Replace '[START]' and '[END]' with your actual target substrings df['extracted_text'] = df['content'].str.extract(r'(?<=\[START\])(.*?)(?=\[END\])', expand=False) # Fill NaN values (from rows where start/end weren't found) with empty string df['extracted_text'] = df['extracted_text'].fillna("") print(df)
Quick Notes:
- If your start/end substrings contain special regex characters (like
.,*,[), escape them usingre.escape()to avoid errors:import re start = re.escape('[START]') end = re.escape('[END]') df['extracted_text'] = df['content'].str.extract(r'(?<=' + start + ')(.*?)(?=' + end + ')', expand=False) - The vectorized method is always preferred for big datasets—it cuts down on processing time significantly.
内容的提问来源于stack exchange,提问作者lalatei

