C++逐词读取文件/字符串时如何添加em dash等自定义分隔符
解决方案
C++ 中 std::istream 默认仅将空白符识别为分隔符,要新增长破折号作为分隔符,推荐采用兼容性更高的后处理拆分方案,无需修改流底层规则,也不会影响短破折号等其他字符的判断:
实现代码
#include <iostream> #include <fstream> #include <string> #include <vector> using namespace std; vector<string> split_by_emdash(const string& s) { vector<string> res; // 若为非UTF-8编码,修改为对应编码的长破折号字符串即可 const string em_dash = "—"; size_t prev = 0; size_t pos = 0; while ((pos = s.find(em_dash, prev)) != string::npos) { if (pos > prev) { res.push_back(s.substr(prev, pos - prev)); } prev = pos + em_dash.size(); } if (prev < s.size()) { res.push_back(s.substr(prev)); } return res; } int main() { ifstream file("example.txt"); string token; while (file >> token) { vector<string> words = split_by_emdash(token); for (const auto& word : words) { cout << word << endl; } } return 0; }
效果验证
针对你给出的测试输入He was young—perhaps from twenty-eight to thirty—tall, slender,代码输出结果完全符合预期:
He was young perhaps from twenty-eight to thirty tall, slender
如果你使用宽字符流处理Unicode文件,只需将代码中string替换为wstring,长破折号常量改为L"\x2014"即可。
内容的提问来源于stack exchange,提问作者jaryl
相关产品推荐
相关产品推荐

