基于Python TF-IDF算法:为词汇加归属标签并过滤特定文档词汇
Absolutely! I've put together a step-by-step Python script that handles everything you asked for: calculating term frequencies across your documents, adding a tag to identify if a term only exists in book3 or appears across all docs, and exporting both full and filtered CSV files (with book3-only terms removed). Let's dive in:
First, make sure you have these packages installed (they'll handle text processing and CSV generation):
pip install pandas scikit-learn
Assuming your documents are saved as text files (e.g., book1.txt, book2.txt, book3.txt), we'll read them into a list first:
import pandas as pd from sklearn.feature_extraction.text import CountVectorizer # Define your document paths and names doc_paths = ["book1.txt", "book2.txt", "book3.txt"] doc_names = ["book1", "book2", "book3"] # Load each document into memory documents = [] for path in doc_paths: with open(path, 'r', encoding='utf-8') as file: documents.append(file.read())
We'll use CountVectorizer to count how many times each term appears in every document, then add a custom tag to flag terms that only exist in book3:
# Initialize vectorizer (set stop_words='english' to remove common words like "the" or "and") vectorizer = CountVectorizer(stop_words='english') term_counts_per_doc = vectorizer.fit_transform(documents) all_terms = vectorizer.get_feature_names_out() # Convert results to a DataFrame for easy manipulation term_df = pd.DataFrame( term_counts_per_doc.toarray().T, # Transpose to have terms as rows index=all_terms, columns=doc_names ) # Calculate total frequency across all documents term_df['total_frequency'] = term_df.sum(axis=1) # Add the presence tag column def determine_presence(row): if row['book1'] == 0 and row['book2'] == 0 and row['book3'] > 0: return "仅出现在book3" elif row['book1'] > 0 and row['book2'] > 0 and row['book3'] > 0: return "出现在所有文档" else: return "出现在部分文档" # Covers cases like terms in book1+book2 only term_df['归属标签'] = term_df.apply(determine_presence, axis=1)
First, let's save a complete CSV with all terms and their tags for reference:
term_df.to_csv('all_terms_with_tags.csv', encoding='utf-8-sig') print("Full CSV with all terms and tags saved as 'all_terms_with_tags.csv'")
Finally, we'll create a filtered CSV that removes any terms marked as "仅出现在book3":
# Filter out book3-only terms filtered_terms_df = term_df[term_df['归属标签'] != "仅出现在book3"] # Export the filtered results filtered_terms_df.to_csv('filtered_terms.csv', encoding='utf-8-sig') print("Filtered CSV (without book3-only terms) saved as 'filtered_terms.csv'")
- If you don't want to remove stopwords, change
CountVectorizer(stop_words='english')toCountVectorizer(). - If you need TF-IDF scores instead of raw term counts, replace
CountVectorizerwithTfidfVectorizer(fromsklearn.feature_extraction.text). - Adjust the tag labels (e.g., use English like "Only in book3" or "Present in all docs") if needed.
- For non-txt documents (like .docx), use libraries like
python-docxto read content instead of plain file open.
内容的提问来源于stack exchange,提问作者Camilla8

