如何用C++正则表达式提取文件中精确10位长度的数字?
Hey there! I get exactly what you're dealing with—your current regex is grabbing any 10-digit chunk out of longer numbers, but you want to target only independent 10-digit numbers that aren't part of a longer numeric string. Let's sort this out quickly.
The Problem with Your Current Regex
Your regex [0-9]{10} (or shorthand \d{10}) just looks for 10 consecutive digits, regardless of what comes before or after them. That's why it's pulling "0987654321" from the 12-digit string "098765432123"—it doesn't care that extra digits are attached to the sequence.
Two Solutions to Match Standalone Numbers
We need to tell the regex that the 10-digit number shouldn't be surrounded by other digits. Here are two reliable approaches:
1. Word Boundaries (\b)
Word boundaries (\b) match the edge between a "word character" (letters, digits, underscores) and a non-word character (spaces, punctuation, start/end of string). For most common cases (like numbers separated by spaces or punctuation), this works perfectly.
Use this regex (using a raw string literal to avoid escaping headaches in C++):
std::regex re(R"(\b\d{10}\b)");
2. Negative Lookaround Assertions
If you need something more strict (e.g., numbers might be next to underscores or other symbols that count as word characters), use negative lookbehind and lookahead. These check that there's no digit immediately before or after the 10-digit sequence:
std::regex re(R"((?<!\d)\d{10}(?!\d))");
(?<!\d): Ensures the sequence is not preceded by a digit(?!\d): Ensures the sequence is not followed by a digit
Modified Full Code
Here's your updated code with both options (uncomment the one that fits your use case):
#include <fstream> #include <iostream> #include <regex> #include <string> int main() { std::string subject("098765432123 1234567890 12345 67890123456 _1234567890_"); try { // Option 1: Word boundaries (great for standard space/punctuation-separated numbers) std::regex re(R"(\b\d{10}\b)"); // Option 2: Negative lookaround (more strict, handles symbols like underscores) // std::regex re(R"((?<!\d)\d{10}(?!\d))"); std::sregex_iterator next(subject.begin(), subject.end(), re); std::sregex_iterator end; while (next != end) { std::smatch match = *next; std::cout << "Matched: " << match.str() << "\n"; next++; } } catch (const std::regex_error& e) { std::cerr << "Regex error: " << e.what() << "\n"; } return 0; }
Test Results
- With Option 1, the output will be just
Matched: 1234567890(the_1234567890_won't match because underscores are considered word characters). - With Option 2, the output will include both
1234567890and1234567890(from the underscore example), since underscores aren't digits.
Pick the option that fits your specific needs—both will fix your original problem of matching substrings from longer numbers!
内容的提问来源于stack exchange,提问作者Benjamin Sx

