如何将istream与正则表达式结合处理HTML文件并替换标签内容?
Using
istream with Regular Expressions to Modify HTML Table Content Great question! When dealing with large HTML files, streaming content (instead of loading everything into memory) is the way to go—and combining std::istream with C++'s standard regex library is totally feasible. Let’s break this down with a practical example and step-by-step explanations.
Key Concepts to Know
First, C++11 and later include the <regex> header, which works seamlessly with istream by letting you process content as you read it. We’ll use:
std::ifstreamto read the HTML file streamstd::regexto define patterns for<tr>and<td>tagsstd::sregex_iteratorto iterate over matches in each chunk of the stream- Non-greedy matching (
.*?) to avoid accidentally capturing too much content (critical for nested HTML structures)
Complete Example Code
This code will locate <tr> tags, find their child <td> elements, and replace the <td> content with your custom values:
#include <fstream> #include <iostream> #include <regex> #include <string> #include <cerrno> using namespace std; int main() { // Open the input HTML file ifstream infile("html9.html"); if (!infile.is_open()) { cerr << "Error opening file: " << strerror(errno) << endl; return 1; } // Prepare output file (optional, replace cout if you want to save changes) ofstream outfile("modified_html9.html"); if (!outfile.is_open()) { cerr << "Error creating output file: " << strerror(errno) << endl; infile.close(); return 1; } // Regex patterns: // Match <tr> tags (with optional attributes) and their content (non-greedy) regex tr_pattern(R"(<tr\b[^>]*>(.*?)</tr>)", regex::icase | regex::dotall); // Match <td> tags and capture their inner content regex td_pattern(R"(<td\s*>(.*?)</td>)", regex::icase | regex::dotall); string line; // Read the file line by line (streaming, no full file in memory) while (getline(infile, line)) { string processed_line = line; // Find all <tr> blocks in the current line sregex_iterator tr_it(processed_line.begin(), processed_line.end(), tr_pattern); sregex_iterator tr_end; for (; tr_it != tr_end; ++tr_it) { string original_tr = (*tr_it)[0]; // Full <tr> block including tags string tr_content = (*tr_it)[1]; // Content inside the <tr> tags // Modify the <td> elements within this <tr> block string modified_tr_content = regex_replace( tr_content, td_pattern, [](const smatch& match) { string td_inner = match[1].str(); // Your custom replacement logic here if (td_inner == " #Title ") { return "<td> My Custom Title </td>"; } else if (td_inner == " #Info ") { return "<td> My Custom Info </td>"; } else { // Keep original <td> if no match return match[0].str(); } } ); // Replace the original <tr> block with the modified version size_t tr_pos = processed_line.find(original_tr); processed_line.replace(tr_pos, original_tr.size(), "<tr>" + modified_tr_content + "</tr>"); } // Write the processed line to output (or print to console) outfile << processed_line << endl; // cout << processed_line << endl; // Uncomment to print to console } // Cleanup infile.close(); outfile.close(); cout << "Modification complete! Check modified_html9.html" << endl; return 0; }
Step-by-Step Explanation
1. Setup & File Handling
- We open both an input (
ifstream) and optional output (ofstream) file. Always check if files opened successfully—this avoids silent failures. - Using
getlinereads one line at a time, which keeps memory usage low even for huge HTML files.
2. Regex Pattern Design
<tr>Pattern:R"(<tr\b[^>]*>(.*?)</tr>)"\bensures we match the whole word "tr" (not part of another tag like<strong>)[^>]*matches any attributes inside the<tr>tag (e.g.,<tr class="highlight">).*?is non-greedy, so it stops at the first</tr>instead of matching all the way to the last one in the fileregex::icaseignores case (matches<TR>,<Tr>, etc.),regex::dotalllets.match newline characters (for multi-line<tr>blocks)
<td>Pattern:R"(<td\s*>(.*?)</td>)"\s*matches any spaces after<td>- Captures the inner content of the
<td>with(.*?)so we can replace it
3. Processing Each Line
- For each line, we use
sregex_iteratorto find all<tr>blocks. This iterator walks through all matches of our<tr>pattern. - For each
<tr>block:- We extract the content inside the
<tr>tags - Use
regex_replacewith a lambda function to customize the replacement logic. The lambda checks the inner content of each<td>and replaces it with your desired text. - We replace the original
<tr>block in the line with the modified version.
- We extract the content inside the
4. Output
- The processed line is written to an output file (or printed to the console if you prefer). This way, you don’t modify the original file directly—always a safe practice!
Notes for Edge Cases
- If your
<tr>blocks span multiple lines, thegetlineapproach might split them. To handle this, you could read chunks of the file instead of lines, or adjust the regex to account for line breaks (whichregex::dotallhelps with). - For more complex HTML, consider using an HTML parser (like Gumbo or libxml2) instead of regex—regex isn’t perfect for nested HTML structures, but it works great for simple, predictable tables like yours.
内容的提问来源于stack exchange,提问作者user786
相关产品推荐
相关产品推荐

