如何用Linux sed命令去除首尾文本,保留中间可变内容并格式化输出?
Solution for Extracting Text with Sed from HTML Line
Let's break down how to get your desired output using sed. The goal is to strip away the surrounding HTML tags and prepend Title: to the variable text inside.
Basic Sed Command
Here's a straightforward command that works for your specific HTML line:
echo '<div class="u-flL condText " id="vi-itm-cond" itemprop="itemCondition"> Seller refurbished </div>' | sed -E 's/^<div[^>]*>(.*)<\/div>$/Title: \1/'
What This Does:
^<div[^>]*>: Matches the opening<div>tag and all its attributes (the[^>]*part captures everything until the closing>of the tag, so it works even if the div has different classes/IDs).(.*): A capture group that grabs all the text between the opening and closing tags (your variable "Seller refurbished" content, including any leading/trailing spaces).<\/div>$: Matches the closing</div>tag at the end of the line (the backslash escapes the forward slash so sed recognizes it as part of the tag).Title: \1: Replaces the entire line withTitle:followed by the text captured in the first group.
Trim Leading/Trailing Spaces (Optional)
If you want to clean up any extra spaces around your variable text, adjust the regex to target whitespace explicitly:
echo '<div class="u-flL condText " id="vi-itm-cond" itemprop="itemCondition"> Seller refurbished </div>' | sed -E 's/^<div[^>]*>\s*(.*?)\s*<\/div>$/Title: \1/'
\s*: Matches any number of whitespace characters (spaces, tabs, etc.) before and after your target text..*?: Uses non-greedy matching to stop as soon as it hits the closing whitespace, ensuring we only capture the actual content without extra padding.
Testing the Command
When you run either of these commands with your sample HTML line, you'll get exactly what you need:
Title: Seller refurbished
Quick Notes:
- This assumes your HTML is on a single line (like your example). If the HTML spans multiple lines, sed alone might not be ideal—you'd need to first join lines or use a tool like
awkthat handles multi-line input better. - The regex is tailored to match a single
<div>element wrapping your text. If there are nested tags inside the div, this will capture all content between the opening and closing div tags (adjust the regex if you need to target more specific nested content).
内容的提问来源于stack exchange,提问作者Sergey Ponitaiev
相关产品推荐
相关产品推荐

