如何提取文本文件中多次出现指定词汇后的3000字符或300词?
How to Extract 3000 Characters or 300 Words After Each "Accounting Principles" Instance
Got it, let's walk through how to pull the text following each occurrence of Accounting Principles—capped at either 3000 characters or 300 words, whichever comes first. Here are two reliable methods:
1. Python Script (Automated, Scalable)
This is the best option if you're working with longer texts or need to repeat this task. The script will find every hit and extract the desired text automatically:
def extract_post_keyword_text(input_text, keyword, max_chars=3000, max_words=300): extracted_sections = [] current_pos = 0 while True: # Locate the next occurrence of the keyword keyword_pos = input_text.find(keyword, current_pos) if keyword_pos == -1: break # No more occurrences to find # Start extracting right after the keyword ends start_extract = keyword_pos + len(keyword) # First grab up to 3000 characters char_limit_text = input_text[start_extract : start_extract + max_chars] # Now trim to 300 words if that's shorter word_list = char_limit_text.split() word_limit_text = ' '.join(word_list[:max_words]) # Add the section to results (include context for clarity) extracted_sections.append(f""" ### Occurrence at index {keyword_pos}: > **{keyword}** {word_limit_text} """) # Move the search position forward to avoid reprocessing the same text current_pos = start_extract return extracted_sections # Your input text source_text = """Accounting Principles. Negative Pledge Clauses . Clauses Restricting Subsidiary Distributions . Lines of Business......Accounting Principles: is defined in the definition of IFRS. Administrative Agent: SVB......In the event that any Accounting Principles (as defined below) shall occur and such change results......""" # Run the extraction results = extract_post_keyword_text(source_text, "Accounting Principles") # Print or save the output for section in results: print(section.strip())
What this script does:
- Loops through every instance of
Accounting Principlesin your text. - For each hit, it first extracts up to 3000 characters starting right after the keyword.
- Then it splits that text into words and caps it at 300 words (so you don't end up with partial words or overly long sections if 300 words is shorter than 3000 characters).
- Outputs each section with a clear label showing where the occurrence was found in the original text.
2. Manual Method (Quick, No Code Needed)
If you just need a one-off extraction, use a text editor like VS Code, Sublime Text, or even Notepad++:
- Use the Find All function (press
Ctrl+Shift+Fon Windows/Linux,Cmd+Shift+Fon Mac) and search forAccounting Principles. - For each result in the find panel, click to jump to that position in the text.
- Place your cursor right after the keyword, then:
- To count characters: Watch the character counter in your editor's status bar while selecting text until you hit 3000.
- To count words: Hold
Ctrl(orCmd) and click repeatedly to select individual words until you reach 300, then copy the selection.
Either method will get you the text you need. The script is great for consistency or larger datasets, while the manual approach works for quick checks.
内容的提问来源于stack exchange,提问作者WUTONG
相关产品推荐
相关产品推荐

