You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取字符串中可预测标记间的包含式文本?求更优实现方案

Cleaner Approach to Extract Sections Between Predictable [http:*] Markers

Hey there! I see you've got a working solution but it's a bit messy and tough to maintain down the line. Let's swap that out for a much cleaner, more maintainable approach using regular expressions—they're made exactly for this kind of pattern-matching task.

The Problem Recap

You want to split a string into sections where each section starts with a [http:...] marker, includes all text until the next [http:...] marker (or the end of the string), and keeps the marker itself.

The Regex Solution

Regex lets us define the pattern we want to match in one place, making the code concise and easy to tweak if your marker format ever changes. Here's the pattern we'll use, plus a breakdown:

\[http:[^\]]+\].*?(?=\[http:|$)
  • \[http:[^\]]+\]: Matches the full [http:Something] marker—this captures everything from [http: up to the closing ]
  • .*?: Non-greedy match for any characters (this ensures we stop at the next [http: marker instead of running all the way to the end of the string)
  • (?=\[http:|$): A positive lookahead that tells the regex to stop when it sees either the next [http: marker or the end of the string

C# Implementation

Here's the full, clean code that uses this regex to get your desired output:

using System;
using System.Collections.Generic;
using System.Text.RegularExpressions;

public class SectionExtractor
{
    public static List<string> ExtractSections(string input)
    {
        List<string> sections = new List<string>();
        // Define our pattern for matching each [http:*] section
        string sectionPattern = @"\[http:[^\]]+\].*?(?=\[http:|$)";
        
        // Find all matches in the input string
        MatchCollection matches = Regex.Matches(input, sectionPattern);
        
        foreach (Match match in matches)
        {
            // Trim any trailing whitespace to match your ideal output exactly
            sections.Add(match.Value.TrimEnd());
        }
        
        return sections;
    }

    public static void Main()
    {
        string fullString = "[http:Something] One Two Three [http:AnotherOne] Four [http:BlahBlah] sdksaod,cne 9ofew {}@:P{";
        List<string> result = ExtractSections(fullString);
        
        // Print the results like your example
        for (int i = 0; i < result.Count; i++)
        {
            Console.WriteLine($"String{i+1} = `{result[i]}`");
        }
    }
}

Why This Is Better Than Your Current Code

  • No fragile loop flags: Your existing code uses found and temp variables which are easy to mess up if you need to adjust logic later
  • Handles edge cases: If your [http:...] markers ever had spaces inside them (unlikely now, but possible), your split-by-space approach would break—this regex doesn't care about spaces in the marker
  • Easier to maintain: If your marker format changes (e.g., switches to [https:...]), you just update one line of the regex instead of rewriting loop logic
  • More readable: Anyone looking at this code can immediately see we're matching a specific pattern, rather than parsing through a loop of split strings

内容的提问来源于stack exchange,提问作者Adam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:39:15