You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将istream与正则表达式结合处理HTML文件并替换标签内容?

Using istream with Regular Expressions to Modify HTML Table Content

Great question! When dealing with large HTML files, streaming content (instead of loading everything into memory) is the way to go—and combining std::istream with C++'s standard regex library is totally feasible. Let’s break this down with a practical example and step-by-step explanations.

Key Concepts to Know

First, C++11 and later include the <regex> header, which works seamlessly with istream by letting you process content as you read it. We’ll use:

  • std::ifstream to read the HTML file stream
  • std::regex to define patterns for <tr> and <td> tags
  • std::sregex_iterator to iterate over matches in each chunk of the stream
  • Non-greedy matching (.*?) to avoid accidentally capturing too much content (critical for nested HTML structures)

Complete Example Code

This code will locate <tr> tags, find their child <td> elements, and replace the <td> content with your custom values:

#include <fstream>
#include <iostream>
#include <regex>
#include <string>
#include <cerrno>

using namespace std;

int main() {
    // Open the input HTML file
    ifstream infile("html9.html");
    if (!infile.is_open()) {
        cerr << "Error opening file: " << strerror(errno) << endl;
        return 1;
    }

    // Prepare output file (optional, replace cout if you want to save changes)
    ofstream outfile("modified_html9.html");
    if (!outfile.is_open()) {
        cerr << "Error creating output file: " << strerror(errno) << endl;
        infile.close();
        return 1;
    }

    // Regex patterns:
    // Match <tr> tags (with optional attributes) and their content (non-greedy)
    regex tr_pattern(R"(<tr\b[^>]*>(.*?)</tr>)", regex::icase | regex::dotall);
    // Match <td> tags and capture their inner content
    regex td_pattern(R"(<td\s*>(.*?)</td>)", regex::icase | regex::dotall);

    string line;
    // Read the file line by line (streaming, no full file in memory)
    while (getline(infile, line)) {
        string processed_line = line;

        // Find all <tr> blocks in the current line
        sregex_iterator tr_it(processed_line.begin(), processed_line.end(), tr_pattern);
        sregex_iterator tr_end;

        for (; tr_it != tr_end; ++tr_it) {
            string original_tr = (*tr_it)[0]; // Full <tr> block including tags
            string tr_content = (*tr_it)[1];  // Content inside the <tr> tags

            // Modify the <td> elements within this <tr> block
            string modified_tr_content = regex_replace(
                tr_content,
                td_pattern,
                [](const smatch& match) {
                    string td_inner = match[1].str();
                    // Your custom replacement logic here
                    if (td_inner == " #Title ") {
                        return "<td> My Custom Title </td>";
                    } else if (td_inner == " #Info ") {
                        return "<td> My Custom Info </td>";
                    } else {
                        // Keep original <td> if no match
                        return match[0].str();
                    }
                }
            );

            // Replace the original <tr> block with the modified version
            size_t tr_pos = processed_line.find(original_tr);
            processed_line.replace(tr_pos, original_tr.size(), "<tr>" + modified_tr_content + "</tr>");
        }

        // Write the processed line to output (or print to console)
        outfile << processed_line << endl;
        // cout << processed_line << endl; // Uncomment to print to console
    }

    // Cleanup
    infile.close();
    outfile.close();
    cout << "Modification complete! Check modified_html9.html" << endl;
    return 0;
}

Step-by-Step Explanation

1. Setup & File Handling

  • We open both an input (ifstream) and optional output (ofstream) file. Always check if files opened successfully—this avoids silent failures.
  • Using getline reads one line at a time, which keeps memory usage low even for huge HTML files.

2. Regex Pattern Design

  • <tr> Pattern: R"(<tr\b[^>]*>(.*?)</tr>)"
    • \b ensures we match the whole word "tr" (not part of another tag like <strong>)
    • [^>]* matches any attributes inside the <tr> tag (e.g., <tr class="highlight">)
    • .*? is non-greedy, so it stops at the first </tr> instead of matching all the way to the last one in the file
    • regex::icase ignores case (matches <TR>, <Tr>, etc.), regex::dotall lets . match newline characters (for multi-line <tr> blocks)
  • <td> Pattern: R"(<td\s*>(.*?)</td>)"
    • \s* matches any spaces after <td>
    • Captures the inner content of the <td> with (.*?) so we can replace it

3. Processing Each Line

  • For each line, we use sregex_iterator to find all <tr> blocks. This iterator walks through all matches of our <tr> pattern.
  • For each <tr> block:
    • We extract the content inside the <tr> tags
    • Use regex_replace with a lambda function to customize the replacement logic. The lambda checks the inner content of each <td> and replaces it with your desired text.
    • We replace the original <tr> block in the line with the modified version.

4. Output

  • The processed line is written to an output file (or printed to the console if you prefer). This way, you don’t modify the original file directly—always a safe practice!

Notes for Edge Cases

  • If your <tr> blocks span multiple lines, the getline approach might split them. To handle this, you could read chunks of the file instead of lines, or adjust the regex to account for line breaks (which regex::dotall helps with).
  • For more complex HTML, consider using an HTML parser (like Gumbo or libxml2) instead of regex—regex isn’t perfect for nested HTML structures, but it works great for simple, predictable tables like yours.

内容的提问来源于stack exchange,提问作者user786

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 12:32:37