You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas提取CSV关键词并匹配另一个CSV内容?

How to Dynamically Extract Keywords from a CSV for Matching with Pandas

Hey there! I totally get why you want to move away from hardcoding keywords—once that list grows, maintaining a long string like Apple|Banana|Orange becomes a total headache. Let's adjust your code to pull keywords directly from CSV1 and use them to match content in CSV2 seamlessly.

Step-by-Step Solution

Here's a revised version of your code that handles dynamic keyword extraction:

import pandas as pd
import re

# 1. Load the keyword CSV and extract the list of keywords
# Replace 'keyword' with the actual column name in your CSV1.csv
keywords_df = pd.read_csv('CSV1.csv')
keywords_list = keywords_df['keyword'].dropna().tolist()  # Drop any empty keyword entries

# 2. Create a regex pattern from the keywords
# Use re.escape() to handle any regex special characters (like ., *, etc.) in keywords
keyword_pattern = '|'.join(re.escape(keyword) for keyword in keywords_list)

# 3. Load the content CSV and filter for matches
# Replace 'content' with the actual column name in your CSV2.csv
content_df = pd.read_csv('CSV2.csv')
matching_results = content_df[content_df['content'].str.contains(keyword_pattern, case=False)]

# View the matched entries
print(matching_results)

Key Details to Note

  • Dynamic Keyword Loading: By using .tolist() on the keyword column, we automatically pull all entries from CSV1—no more updating code every time your keyword list expands.
  • Regex Safety: The re.escape() function ensures that any special characters in your keywords (like Grapefruit! or Pear.*) don't break the regex matching. If you're 100% sure your keywords have no special characters, you can skip this, but it's a safe practice to include.
  • Case Insensitivity: The case=False parameter lets you match entries like "I like Apples" even if your keyword is "apple". Remove this if you need exact case matching.
  • Cleaning Data: The .dropna() call filters out any empty rows in your keyword CSV, which prevents invalid empty matches from cluttering results.

Quick Adjustments for Your Use Case

  • Double-check the column names in both CSVs: If CSV1's keywords are in a column named fruit instead of keyword, update keywords_df['keyword'] to keywords_df['fruit']. Same goes for CSV2's content column—swap content for whatever your actual column name is.

内容的提问来源于stack exchange,提问作者ele

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 09:02:42