如何提取文本文件中数字前的单词?Python代码实现求助
Let’s walk through a practical, Python-based approach to solve your problem—we’ll safely read your text files, filter lines containing your target content, and pull out the words immediately preceding numeric sequences.
Step-by-Step Breakdown
1. Safe File Handling with Context Managers
Instead of using f = open(FILENAME, 'r') directly, we’ll use a with statement. This automatically closes the file after processing, preventing resource leaks and making your code more robust.
2. Filter Lines with Specific Content
First, define the string you want to target (like "Error"), then skip any lines that don’t include this content to focus only on relevant lines.
3. Regular Expressions for Pattern Matching
We’ll use regex to identify words followed by numbers. The pattern r'(\w+)\s+(\d+)' captures any word (letters/underscores) that comes right before whitespace and one or more digits. Using re.findall() lets us grab all such pairs in a single line.
Full Code Example
import re FILENAME = "your_file.txt" # Replace with your actual file path TARGET_CONTENT = "Error" # Replace with your specific content to locate lines def extract_words_before_numbers(): with open(FILENAME, 'r') as f: for line in f: cleaned_line = line.strip() # Skip lines that don't contain our target content if TARGET_CONTENT not in cleaned_line: continue # Find all word-number pairs in the line matches = re.findall(r'(\w+)\s+(\d+)', cleaned_line) # Extract just the words from each matched pair target_words = [match[0] for match in matches] print(f"Words found in line: {target_words}") if __name__ == "__main__": extract_words_before_numbers()
Testing with Your Sample Inputs
Let’s see how this works with your example lines:
Sample 1 Line:
Error you can find the Error in Section Nose 546516547 in Lower Part 211547
Output:['Nose', 'Part']
Sample 2 Line:
Error the file break at you knee 3215467 in the element 51423165
Output:['knee', 'element']
Customization Tips
- If your "words" include special characters (like hyphens or apostrophes), adjust the regex to
r'([\w-]+)\s+(\d+)'to include hyphens, or expand the character set as needed. - To match only specific number lengths (e.g., 7-9 digits), modify the digit part:
r'(\w+)\s+(\d{7,9})'. - Instead of printing results, you can collect them in a list or write them to a new file for later use.
Quick Notes
- Don’t forget to import the
remodule—it’s required for regex operations. - The
strip()method removes leading/trailing whitespace from lines, ensuring cleaner pattern matching.
内容的提问来源于stack exchange,提问作者classic

