使用Python与spaCy提取支出冻结信息的技术指导请求
Hey there! As someone starting out with NLP and spaCy, this is a great practical project to learn on. Let's walk through exactly how to extract those expense freeze statuses from your text and export the results to Excel. I'll keep things straightforward so you can follow along easily.
Step 1: Lay Out Your Target Expense Categories
First, let's list out all the expense types we need to check—this matches the table you provided:
- purchase order
- capital
- consulting
- business meetings
- external hires
- KM&L
- travel
We'll use this list to scan the text for each category and determine its freeze status.
Step 2: Process the Text with spaCy
You already have the basics of loading spaCy and processing the text, so let's build on that. We'll add logic to detect whether each expense category is frozen or not by looking for key phrases like frozen, freeze, not frozen, or not on 'freeze' in the context of each category.
Here's a modified version of your code with added detection logic:
import spacy from spacy.matcher import PhraseMatcher # Load spaCy model nlp = spacy.load('en_core_web_sm') # Your target text (cleaned up for consistency) text = """Non-revenue-generating purchase order expenditures will be frozen. All capital related expenditures are frozen effectively for Q4. Following spending categories are frozen: Consulting, (including existing engagements), Business meetings. Please note that there is a hiring freeze for external hires, subcontractors and consulting services. KM&L expenditure will not be frozen. Travel cost will not be on ‘freeze’.""" # Process the text doc = nlp(text) # Define our expense categories expense_categories = [ "purchase order", "capital", "consulting", "business meetings", "external hires", "KM&L", "travel" ] # Create a list to store results results = [] # Check each category for category in expense_categories: # Initialize statuses frozen = None not_frozen = None # Set up phrase matcher to find exact category matches matcher = PhraseMatcher(nlp.vocab) matcher.add(category, None, nlp(category)) matches = matcher(doc) if matches: # Get the sentence where the category appears match_id, start, end = matches[0] span = doc[start:end] sentence_text = span.sent.text.lower() # Check for freeze status in the sentence if "frozen" in sentence_text or "freeze" in sentence_text: if "not frozen" in sentence_text or "not on ‘freeze’" in sentence_text: not_frozen = "not frozen" else: frozen = "frozen" # Add result to our list results.append({ "TYPE_OF_EXPENSE": category, "FROZEN?": frozen, "NOT_FROZEN?": not_frozen }) # Verify results in console for res in results: print(res)
How This Works:
- We use spaCy's
PhraseMatcherto find exact matches of each expense category—this is more reliable than basic string matching because it accounts for proper tokenization. - For each matched category, we analyze the entire sentence it appears in to check for freeze-related terms.
- We look for negations like "not" paired with freeze terms to distinguish between frozen and unfrozen statuses.
Step 3: Export Results to Excel
To export the results to Excel, we'll use pandas—it's the simplest way to handle tabular data and export to Excel. If you don't have pandas installed, run pip install pandas openpyxl first (openpyxl is required for Excel file generation).
Add this code after the results are generated:
import pandas as pd # Convert results to a DataFrame df = pd.DataFrame(results) # Export to Excel (index=False removes the default row numbers) df.to_excel("expense_freeze_status.xlsx", index=False) print("Results exported to expense_freeze_status.xlsx successfully!")
Expected Output
When you run this code, you'll get an Excel file with exactly the table you wanted:
| TYPE_OF_EXPENSE | FROZEN? | NOT_FROZEN? |
|---|---|---|
| purchase order | frozen | None |
| capital | frozen | None |
| consulting | frozen | None |
| business meetings | frozen | None |
| external hires | frozen | None |
| KM&L | None | not frozen |
| travel | None | not frozen |
Tips for Future Improvement
As you get more comfortable with spaCy, you could:
- Use dependency parsing to better detect negations (instead of just checking for exact phrases like "not frozen"). For example, identifying if a negation token (
not) is grammatically linked to the freeze term. - Train a small custom spaCy pipeline to automatically classify freeze statuses if your text becomes more complex over time.
内容的提问来源于stack exchange,提问作者Radu M

