基于输入删除DataFrame行:Python是否有类SQL LIKE的功能?
Hey there! Great questions—both are totally achievable with pandas (I’m assuming that’s the library you’re using for your DataFrame, since it’s the standard in Python for this kind of work). Let’s break this down step by step:
1. Simulating SQL LIKE Functionality in Pandas
Pandas has a super handy str.contains() method that works just like SQL’s LIKE operator, and it supports regular expressions for more flexibility. Here are the common use cases:
- Match strings containing a substring (equivalent to
LIKE '%substring%'):# Filter rows where 'your_column' contains 'target' filtered_df = df[df['your_column'].str.contains('target')] - Match strings starting with a substring (equivalent to
LIKE 'prefix%'):
Use the^regex anchor to denote the start of a string:filtered_df = df[df['your_column'].str.contains('^prefix')] - Match strings ending with a substring (equivalent to
LIKE '%suffix'):
Use the$regex anchor for the end of a string:filtered_df = df[df['your_column'].str.contains('suffix$')] - Case-insensitive matching:
Addcase=Falseto ignore uppercase/lowercase differences:filtered_df = df[df['your_column'].str.contains('target', case=False)] - Match multiple patterns (equivalent to
LIKE '%foo%' OR LIKE '%bar%'):
Use the pipe|in regex to separate patterns:filtered_df = df[df['your_column'].str.contains('foo|bar')]
2. Deleting Rows Where the First Column is 'CLS'
This is straightforward too—you have a few options depending on your preference:
- Option 1: Boolean Indexing (most intuitive)
Keep only rows where the first column’s value is NOT 'CLS':# iloc[:, 0] selects the first column regardless of its name df = df[df.iloc[:, 0] != 'CLS'] - Option 2: Using
drop()method
First get the indices of rows where the first column is 'CLS', then drop those indices:cls_indices = df[df.iloc[:, 0] == 'CLS'].index df = df.drop(cls_indices) # Or do it in one line: df = df.drop(df[df.iloc[:, 0] == 'CLS'].index) - Option 3: Using
query()method
If you prefer a more SQL-like syntax:# Get the first column's name and use it in the query first_col = df.columns[0] df = df.query(f"{first_col} != 'CLS'") - Note about in-place modification:
If you want to modify the original DataFrame instead of creating a new one, addinplace=Trueto thedrop()method:df.drop(df[df.iloc[:, 0] == 'CLS'].index, inplace=True)
内容的提问来源于stack exchange,提问作者Bystrov_2108
相关产品推荐
相关产品推荐

