基于父DataFrame列内特定字符串筛选生成多张子DataFrame的技术咨询
Got it, let's break this down step by step. From your description, you need to create 3 subsets from your original DataFrame—all retaining the Response column—where each subset is filtered based on whether the Context column (stored as a comma-separated list of tweets) contains specific target strings. I'll cover the two most common formats your Context column might be in, plus some performance tips.
1. If Context is a Native Python List
If your DataFrame's Context column holds actual Python lists (e.g., ["Tweet1", "Tweet3", "Tweet6"]), you can use a simple apply with membership checks to filter rows:
import pandas as pd # Replace these with your actual target strings key1 = "your_first_target_string" key2 = "your_second_target_string" key3 = "your_third_target_string" # Assume your original DataFrame is named `df` # Subset 1: Rows where Context contains key1 df_sub1 = df[df["Context"].apply(lambda x: key1 in x)].copy() # Should include rows for Tweet1, Tweet3, Tweet6, Tweet7, Tweet11 # Subset 2: Rows where Context contains key1 OR key2 df_sub2 = df[df["Context"].apply(lambda x: key1 in x or key2 in x)].copy() # Should include rows for Tweet1, Tweet2, Tweet3, Tweet4, Tweet6, Tweet7, Tweet8, Tweet11, Tweet12 # Subset 3: Rows where Context contains key1, key2, OR key3 df_sub3 = df[df["Context"].apply(lambda x: any(k in x for k in [key1, key2, key3]))].copy() # Should include rows for Tweet1-Tweet4, Tweet5, Tweet6-Tweet9, Tweet11-Tweet12
2. If Context is a Comma-Separated String
If Context is stored as a string (either unquoted like "Tweet1,Tweet3" or quoted like "['Tweet1', 'Tweet3']"), you'll need to parse it into a list first.
2.1 Unquoted Comma-Separated Strings (e.g., "Tweet1,Tweet3,Tweet6")
Use a helper function to split the string into a clean list:
def str_to_list(s): # Split string and strip any extra whitespace return [item.strip() for item in s.split(",")] # Subset 1 df_sub1 = df[df["Context"].apply(lambda x: key1 in str_to_list(x))].copy() # Subset 2 df_sub2 = df[df["Context"].apply(lambda x: key1 in str_to_list(x) or key2 in str_to_list(x))].copy() # Subset 3 (using `any()` for cleaner code) df_sub3 = df[df["Context"].apply(lambda x: any(k in str_to_list(x) for k in [key1, key2, key3]))].copy()
2.2 Quoted List Strings (e.g., "['Tweet1', 'Tweet3']")
Use ast.literal_eval to safely parse the string into a Python list:
import ast def str_to_list(s): try: return ast.literal_eval(s) except (SyntaxError, ValueError): # Handle invalid entries by returning an empty list return [] # Same filtering code as above df_sub1 = df[df["Context"].apply(lambda x: key1 in str_to_list(x))].copy() df_sub2 = df[df["Context"].apply(lambda x: any(k in str_to_list(x) for k in [key1, key2]))].copy() df_sub3 = df[df["Context"].apply(lambda x: any(k in str_to_list(x) for k in [key1, key2, key3]))].copy()
Pro Tips for Better Performance & Reliability
- Speed Up Large DataFrames: If you're working with a huge dataset,
applycan be slow. For string-basedContextcolumns, usestr.containswith regex instead—it's much faster:import re # Escape special characters in keys to avoid regex issues pattern = "|".join(re.escape(k) for k in [key1, key2, key3]) # Subset 1 (single key) df_sub1 = df[df["Context"].str.contains(key1, regex=False)].copy() # Subset 2 (key1 or key2) df_sub2 = df[df["Context"].str.contains(f"{re.escape(key1)}|{re.escape(key2)}", regex=True)].copy() # Subset 3 (any of the three keys) df_sub3 = df[df["Context"].str.contains(pattern, regex=True)].copy() - Avoid Warnings: Always use
.copy()when creating subsets to preventSettingWithCopyWarninglater if you modify the subset data. - Validate Results: After creating subsets, double-check that the correct rows are included by verifying against your expected tweet IDs (e.g.,
print(df_sub1["TweetID"].tolist())if you have aTweetIDcolumn).
内容的提问来源于stack exchange,提问作者CD_NS

