You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python TF-IDF算法:为词汇加归属标签并过滤特定文档词汇

Absolutely! I've put together a step-by-step Python script that handles everything you asked for: calculating term frequencies across your documents, adding a tag to identify if a term only exists in book3 or appears across all docs, and exporting both full and filtered CSV files (with book3-only terms removed). Let's dive in:

Step 1: Install Required Libraries

First, make sure you have these packages installed (they'll handle text processing and CSV generation):

pip install pandas scikit-learn
Step 2: Load Your Documents

Assuming your documents are saved as text files (e.g., book1.txt, book2.txt, book3.txt), we'll read them into a list first:

import pandas as pd
from sklearn.feature_extraction.text import CountVectorizer

# Define your document paths and names
doc_paths = ["book1.txt", "book2.txt", "book3.txt"]
doc_names = ["book1", "book2", "book3"]

# Load each document into memory
documents = []
for path in doc_paths:
    with open(path, 'r', encoding='utf-8') as file:
        documents.append(file.read())
Step 3: Calculate Term Frequencies & Add Presence Tags

We'll use CountVectorizer to count how many times each term appears in every document, then add a custom tag to flag terms that only exist in book3:

# Initialize vectorizer (set stop_words='english' to remove common words like "the" or "and")
vectorizer = CountVectorizer(stop_words='english')
term_counts_per_doc = vectorizer.fit_transform(documents)
all_terms = vectorizer.get_feature_names_out()

# Convert results to a DataFrame for easy manipulation
term_df = pd.DataFrame(
    term_counts_per_doc.toarray().T,  # Transpose to have terms as rows
    index=all_terms,
    columns=doc_names
)

# Calculate total frequency across all documents
term_df['total_frequency'] = term_df.sum(axis=1)

# Add the presence tag column
def determine_presence(row):
    if row['book1'] == 0 and row['book2'] == 0 and row['book3'] > 0:
        return "仅出现在book3"
    elif row['book1'] > 0 and row['book2'] > 0 and row['book3'] > 0:
        return "出现在所有文档"
    else:
        return "出现在部分文档"  # Covers cases like terms in book1+book2 only

term_df['归属标签'] = term_df.apply(determine_presence, axis=1)
Step 4: Export Full CSV (Including All Terms)

First, let's save a complete CSV with all terms and their tags for reference:

term_df.to_csv('all_terms_with_tags.csv', encoding='utf-8-sig')
print("Full CSV with all terms and tags saved as 'all_terms_with_tags.csv'")
Step 5: Filter Out Book3-Only Terms & Export Final CSV

Finally, we'll create a filtered CSV that removes any terms marked as "仅出现在book3":

# Filter out book3-only terms
filtered_terms_df = term_df[term_df['归属标签'] != "仅出现在book3"]

# Export the filtered results
filtered_terms_df.to_csv('filtered_terms.csv', encoding='utf-8-sig')
print("Filtered CSV (without book3-only terms) saved as 'filtered_terms.csv'")
Quick Customization Tips
  • If you don't want to remove stopwords, change CountVectorizer(stop_words='english') to CountVectorizer().
  • If you need TF-IDF scores instead of raw term counts, replace CountVectorizer with TfidfVectorizer (from sklearn.feature_extraction.text).
  • Adjust the tag labels (e.g., use English like "Only in book3" or "Present in all docs") if needed.
  • For non-txt documents (like .docx), use libraries like python-docx to read content instead of plain file open.

内容的提问来源于stack exchange,提问作者Camilla8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 10:08:47