如何用Pandas提取CSV关键词并匹配另一个CSV内容?
How to Dynamically Extract Keywords from a CSV for Matching with Pandas
Hey there! I totally get why you want to move away from hardcoding keywords—once that list grows, maintaining a long string like Apple|Banana|Orange becomes a total headache. Let's adjust your code to pull keywords directly from CSV1 and use them to match content in CSV2 seamlessly.
Step-by-Step Solution
Here's a revised version of your code that handles dynamic keyword extraction:
import pandas as pd import re # 1. Load the keyword CSV and extract the list of keywords # Replace 'keyword' with the actual column name in your CSV1.csv keywords_df = pd.read_csv('CSV1.csv') keywords_list = keywords_df['keyword'].dropna().tolist() # Drop any empty keyword entries # 2. Create a regex pattern from the keywords # Use re.escape() to handle any regex special characters (like ., *, etc.) in keywords keyword_pattern = '|'.join(re.escape(keyword) for keyword in keywords_list) # 3. Load the content CSV and filter for matches # Replace 'content' with the actual column name in your CSV2.csv content_df = pd.read_csv('CSV2.csv') matching_results = content_df[content_df['content'].str.contains(keyword_pattern, case=False)] # View the matched entries print(matching_results)
Key Details to Note
- Dynamic Keyword Loading: By using
.tolist()on the keyword column, we automatically pull all entries from CSV1—no more updating code every time your keyword list expands. - Regex Safety: The
re.escape()function ensures that any special characters in your keywords (likeGrapefruit!orPear.*) don't break the regex matching. If you're 100% sure your keywords have no special characters, you can skip this, but it's a safe practice to include. - Case Insensitivity: The
case=Falseparameter lets you match entries like "I like Apples" even if your keyword is "apple". Remove this if you need exact case matching. - Cleaning Data: The
.dropna()call filters out any empty rows in your keyword CSV, which prevents invalid empty matches from cluttering results.
Quick Adjustments for Your Use Case
- Double-check the column names in both CSVs: If CSV1's keywords are in a column named
fruitinstead ofkeyword, updatekeywords_df['keyword']tokeywords_df['fruit']. Same goes for CSV2's content column—swapcontentfor whatever your actual column name is.
内容的提问来源于stack exchange,提问作者ele
相关产品推荐
相关产品推荐

