You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.6处理HTML文本文件:仅保留表格相关标签的实现方法

Solution for Cleaning HTML to Keep Only Table Tags

Hey there! Let's break this down step by step since you're new to Python—no worries, this is totally manageable.

Your goal is to process an HTML text file, keep only the table-related tags (<table>, <tr>, <td>, and their closing counterparts), and strip out all other tags (meaning we remove everything from < up to > for non-table tags, leaving just the inner content).

Approach Using Regular Expressions (Beginner-Friendly)

We'll use Python's built-in re module to find and modify tags easily. This is simpler than writing manual string parsing from scratch.

Full Code

import re

def process_html_file(input_file_path, output_file_path):
    # Define the core table-related tag names we want to preserve
    keep_tags = {'table', 'tr', 'td'}
    
    def replace_tag(match):
        # Extract the tag name (e.g., 'td' from '<td class="foo">' or '</td>')
        tag = match.group(1)
        clean_tag = tag.strip('/')  # Remove the slash for closing tags
        
        # Keep the tag if it's in our allowed list, else delete it
        if clean_tag in keep_tags:
            return match.group(0)
        else:
            return ''
    
    # Regex pattern to match any standard HTML tag
    tag_pattern = re.compile(r'<(/?\w+)[^>]*>')
    
    # Open input and output files
    with open(input_file_path, 'r', encoding='utf-8') as infile, open(output_file_path, 'w', encoding='utf-8') as outfile:
        for line in infile:
            # Process each line to remove non-table tags
            processed_line = tag_pattern.sub(replace_tag, line)
            outfile.write(processed_line)

# Example usage (replace these paths with your actual files!)
process_html_file('your_input_file.txt', 'cleaned_output.txt')

Step-by-Step Explanation

  • Import re: This module lets us use regular expressions to find all HTML tags in each line without writing messy manual checks.
  • keep_tags Set: We list the tag names we want to keep (no </> or slashes needed—this works for both opening tags like <tr> and closing tags like </tr>).
  • replace_tag Helper Function: This function runs every time the regex finds a tag. It checks if the tag is in our allowed list. If yes, it keeps the full tag; if not, it replaces the tag with an empty string (deleting it entirely).
  • Regex Pattern: r'<(/?\w+)[^>]*>' matches any standard HTML tag:
    • < = starts with opening angle bracket
    • /? = optional slash (for closing tags)
    • \w+ = the tag name (letters/numbers/underscores)
    • [^>]* = any characters until the closing angle bracket
  • File Handling: We use the with statement to open files safely—it automatically closes them when we're done, so you don't have to worry about forgetting that step. We read each line, process it, and write the cleaned version to the output file.

Tips for Beginners

  • Replace 'your_input_file.txt' and 'cleaned_output.txt' with the actual paths to your files. If the files are in the same folder as your Python script, just use the filename.
  • Test with a small sample file first to make sure it behaves as expected. For example, if your input line is <p>Hello</p><tr><td>Table content</td></tr>, the output will be Hello<tr><td>Table content</td></tr>.
  • If your HTML has unusual tags (like those with special characters in the name), you might need to tweak the regex, but this works for most standard table-based HTML.

内容的提问来源于stack exchange,提问作者statsguy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:01:54