Python正则优化:提取含指定文件名的注释间HTML行
Fixing Regex to Target Specific HTML Block with Filename
Got it, let's sort out that overzealous regex for you! The issue is almost certainly greedy matching causing your pattern to gobble up most of the document—here's how to fix it and target exactly the block containing your specific filename, using Python.
The Solution: Non-Greedy Matching + Precise Targeting
We need a regex that:
- Matches the
<!--Copy from here-->comment (and its case/space variants) - Grabs content only until the next matching comment (non-greedy)
- Ensures the block contains your target filename (e.g.,
SN_2100-22.htm)
Python Code Implementation
import re # Your HTML content (replace with your actual input) html_content = """ <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-24.htm">2100-24</A></TD> <TD>2100 Bioanalyzer - peak find problem when using new Ambion RNA ladder Cat. #7152</TD> <TD>11/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-23.htm">2100-23</A></TD> <TD>2100 Bioanalyzer communication problems when both Biosizing and 2100 Expert SW are active</TD> <TD>10/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-22.htm">2100-22</A></TD> <TD>Incompatibility of 2100 Expert and Microsoft Windows XP Service Pack 2</TD> <TD>09/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-21.htm">2100-21</A></TD> <TD>2100 Bioanalyzer - DNA LabChip Kits and detergent containing PCR buffer</TD> <TD>04/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-20.htm">2100-20</A></TD> <TD>General PC system and settings requirements for 2100 expert software</TD> <TD>04/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD>2100-19</A></TD> <TD>not used</TD> <TD>01/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> <TR valign=top> <TD><A href="SN_2100-18.htm">2100-18</A></TD> <TD>RNA 6000 Pico Kits - ART® Aerosol Resistant Tips generate baseline abnormalities</TD> <TD >01/04</TD><td valign=top><p align="center">I</p></td></tr> <!--Copy from here--> """ # Replace with your target filename target_filename = "SN_2100-22.htm" # Build the regex pattern pattern = re.compile( r'<!--\s*Copy from here\s*-->(.*?{}.*?)<!--\s*Copy from here\s*-->'.format(re.escape(target_filename)), re.DOTALL | re.IGNORECASE ) # Search for the matching block match = pattern.search(html_content) if match: # Extract and clean up the content extracted_block = match.group(1).strip() print("Extracted Block:\n", extracted_block) else: print(f"No block found containing {target_filename}")
Key Fixes & Explanations
- Non-Greedy Matching: Using
.*?instead of.*ensures the regex stops at the first subsequent<!--Copy from here-->comment, instead of greedily matching to the last one. - Comment Variants:
\s*allows for optional spaces around the comment text, andre.IGNORECASEhandles case differences like<!--COPY FROM HERE-->or<!-- copy from here -->. - Safe Filename Handling:
re.escape(target_filename)ensures any special characters in your filename (like-) don't break the regex.
Output for Your Example
Running the code with target_filename = "SN_2100-22.htm" will output:
<TR valign=top> <TD><A href="SN_2100-22.htm">2100-22</A></TD> <TD>Incompatibility of 2100 Expert and Microsoft Windows XP Service Pack 2</TD> <TD>09/04</TD><td valign=top><p align="center">I</p></td></tr>
Extra Notes
- To match all blocks containing the filename (if it appears multiple times), replace
pattern.search()withpattern.findall()to get a list of matching blocks. - If your comment variants are more diverse (e.g.,
<!--Start Copy-->), adjust the regex's comment portion to account for those (e.g.,<!--\s*(Copy from here|Start Copy)\s*-->).
内容的提问来源于stack exchange,提问作者David J.
相关产品推荐
相关产品推荐

