C++中使用多分隔符分割字符串的实现疑问
Got it, let's fix your tokenization issue! You need to split your input string into clean tokens, keeping only hyphens while stripping all other punctuation. Here are two solid ways to implement your tokenize function in C++:
Method 1: Iterative Character Processing (Follows Your Delimiter Parameter)
This approach directly uses the delimiter string you pass in, skipping any delimiter except hyphens to build valid tokens:
#include <string> #include <vector> #include <cctype> void tokenize(const std::string& input, const std::string& delimiters, std::vector<std::string>& tokens) { tokens.clear(); // Reset the output vector first std::string current_token; for (char c : input) { // Check if the character is a delimiter AND not a hyphen bool is_target_delimiter = (delimiters.find(c) != std::string::npos) && (c != '-'); if (!is_target_delimiter) { // Add non-delimiter characters to the current token current_token += c; } else { // If we have a non-empty token, add it to the list and reset if (!current_token.empty()) { tokens.push_back(current_token); current_token.clear(); } } } // Don't forget to add the final token if it's non-empty if (!current_token.empty()) { tokens.push_back(current_token); } }
How It Works:
- We loop through each character in the input string
- Characters that are in your
delimiterslist but aren't hyphens trigger a token split - Hyphens are treated as valid characters and stay part of the token (like in
COSC-1436) - After processing all characters, we add any remaining token to the vector
Method 2: Regex-Based Extraction (Simpler for Your Specific Use Case)
If you don't need to adjust delimiters dynamically, this method directly extracts all sequences of allowed characters (letters, digits, hyphens) which perfectly matches your requirement:
#include <string> #include <vector> #include <regex> void tokenize(const std::string& input, const std::string& /* delimiters */, std::vector<std::string>& tokens) { tokens.clear(); // Regex pattern to match one or more letters, digits, or hyphens std::regex token_pattern(R"([A-Za-z0-9-]+)"); std::sregex_iterator it(input.begin(), input.end(), token_pattern); std::sregex_iterator end; // Iterate through all matches and add to tokens for (; it != end; ++it) { tokens.push_back(it->str()); } }
How It Works:
- The regex
[A-Za-z0-9-]+finds every continuous sequence of valid characters - We use
sregex_iteratorto loop through all matches and collect them as tokens - This automatically ignores all punctuation (like commas, exclamation points, colons) since they don't match the pattern
Testing Your Input
For your input string "Hello, everyone! This is: COSC-1436, SP18", both methods will produce the tokens:["Hello", "everyone", "This", "is", "COSC-1436", "SP18"]
Just call your function as you already do:
std::string code = "Hello, everyone! This is: COSC-1436, SP18"; std::vector<std::string> tokens; tokenize(code, " .,:;!?", tokens);
内容的提问来源于stack exchange,提问作者Tristan

