Python3.6处理HTML文本文件:仅保留表格相关标签的实现方法
Hey there! Let's break this down step by step since you're new to Python—no worries, this is totally manageable.
Your goal is to process an HTML text file, keep only the table-related tags (<table>, <tr>, <td>, and their closing counterparts), and strip out all other tags (meaning we remove everything from < up to > for non-table tags, leaving just the inner content).
Approach Using Regular Expressions (Beginner-Friendly)
We'll use Python's built-in re module to find and modify tags easily. This is simpler than writing manual string parsing from scratch.
Full Code
import re def process_html_file(input_file_path, output_file_path): # Define the core table-related tag names we want to preserve keep_tags = {'table', 'tr', 'td'} def replace_tag(match): # Extract the tag name (e.g., 'td' from '<td class="foo">' or '</td>') tag = match.group(1) clean_tag = tag.strip('/') # Remove the slash for closing tags # Keep the tag if it's in our allowed list, else delete it if clean_tag in keep_tags: return match.group(0) else: return '' # Regex pattern to match any standard HTML tag tag_pattern = re.compile(r'<(/?\w+)[^>]*>') # Open input and output files with open(input_file_path, 'r', encoding='utf-8') as infile, open(output_file_path, 'w', encoding='utf-8') as outfile: for line in infile: # Process each line to remove non-table tags processed_line = tag_pattern.sub(replace_tag, line) outfile.write(processed_line) # Example usage (replace these paths with your actual files!) process_html_file('your_input_file.txt', 'cleaned_output.txt')
Step-by-Step Explanation
- Import
re: This module lets us use regular expressions to find all HTML tags in each line without writing messy manual checks. keep_tagsSet: We list the tag names we want to keep (no</>or slashes needed—this works for both opening tags like<tr>and closing tags like</tr>).replace_tagHelper Function: This function runs every time the regex finds a tag. It checks if the tag is in our allowed list. If yes, it keeps the full tag; if not, it replaces the tag with an empty string (deleting it entirely).- Regex Pattern:
r'<(/?\w+)[^>]*>'matches any standard HTML tag:<= starts with opening angle bracket/?= optional slash (for closing tags)\w+= the tag name (letters/numbers/underscores)[^>]*= any characters until the closing angle bracket
- File Handling: We use the
withstatement to open files safely—it automatically closes them when we're done, so you don't have to worry about forgetting that step. We read each line, process it, and write the cleaned version to the output file.
Tips for Beginners
- Replace
'your_input_file.txt'and'cleaned_output.txt'with the actual paths to your files. If the files are in the same folder as your Python script, just use the filename. - Test with a small sample file first to make sure it behaves as expected. For example, if your input line is
<p>Hello</p><tr><td>Table content</td></tr>, the output will beHello<tr><td>Table content</td></tr>. - If your HTML has unusual tags (like those with special characters in the name), you might need to tweak the regex, but this works for most standard table-based HTML.
内容的提问来源于stack exchange,提问作者statsguy
相关产品推荐
相关产品推荐

