如何在文本文件中提取Consignee详情?需排除Document No等无关内容
Hey there! Let's tackle this issue you're having with extracting only the Consignee details without pulling in Document No and Export Reference lines. Here are a few practical, actionable approaches depending on how your input text is structured:
1. Track Sections & Exclude Unwanted Lines
If you're processing the text line by line, you can flag when you're in the Consignee section, then explicitly skip any lines that mention the other two details. Here's a Python example to illustrate:
consignee_details = [] document_no_details = [] export_ref_details = [] with open("your_input_file.txt", "r") as file: current_section = None for line in file: cleaned_line = line.strip() if not cleaned_line: continue # Skip empty lines to avoid noise # Switch sections when we hit a header if "Consignee" in cleaned_line: current_section = "consignee" continue elif "Document No" in cleaned_line: current_section = "document_no" continue elif "Export Reference" in cleaned_line: current_section = "export_ref" continue # Only add lines to Consignee if they don't belong to the other sections if current_section == "consignee": if "Document No" not in cleaned_line and "Export Reference" not in cleaned_line: consignee_details.append(cleaned_line) elif current_section == "document_no": document_no_details.append(cleaned_line) elif current_section == "export_ref": export_ref_details.append(cleaned_line)
This method keeps you in control of which lines go to which section, with a hard check to block unwanted content from the Consignee list.
2. Define Clear Section Boundaries
If your text uses consistent separators (like empty lines or section headers) between each detail block, you can stop collecting Consignee data as soon as you hit the next section's header. This is even more reliable:
consignee_details = [] collecting_consignee = False with open("your_input_file.txt", "r") as file: for line in file: cleaned_line = line.strip() # Start collecting when we see the Consignee header if "Consignee" in cleaned_line: collecting_consignee = True continue # Stop collecting the moment we hit another section's header if collecting_consignee and ("Document No" in cleaned_line or "Export Reference" in cleaned_line): collecting_consignee = False continue # Add valid lines to the Consignee list if collecting_consignee and cleaned_line: consignee_details.append(cleaned_line)
By setting a hard stop at the next section, you eliminate any chance of accidentally pulling in unrelated lines.
3. Regex with Negative Lookaheads (For Structured Text)
If your text has a predictable format, you can use a regex pattern to match only Consignee lines that don't contain the other two keywords. Here's how that might look:
import re # Regex to match lines that DON'T include Document No or Export Reference valid_consignee_line = re.compile(r'^(?!.*(Document No|Export Reference)).*$') consignee_details = [] collecting_consignee = False with open("your_input_file.txt", "r") as file: for line in file: cleaned_line = line.strip() if "Consignee" in cleaned_line: collecting_consignee = True continue if collecting_consignee and ("Document No" in cleaned_line or "Export Reference" in cleaned_line): collecting_consignee = False continue # Only add lines that pass the regex check if collecting_consignee and valid_consignee_line.match(cleaned_line): consignee_details.append(cleaned_line)
This is great for cases where lines might have partial matches or you need extra precision.
The core idea here is either explicitly filtering out unwanted lines or strictly defining the start/end of the Consignee section. Pick the approach that best fits how consistent your input text's structure is!
内容的提问来源于stack exchange,提问作者Krish

