You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++逐词读取文件/字符串时如何添加em dash等自定义分隔符

解决方案

C++ 中 std::istream 默认仅将空白符识别为分隔符,要新增长破折号作为分隔符,推荐采用兼容性更高的后处理拆分方案,无需修改流底层规则,也不会影响短破折号等其他字符的判断:

实现代码

#include <iostream>
#include <fstream>
#include <string>
#include <vector>

using namespace std;

vector<string> split_by_emdash(const string& s) {
    vector<string> res;
    // 若为非UTF-8编码,修改为对应编码的长破折号字符串即可
    const string em_dash = "—";
    size_t prev = 0;
    size_t pos = 0;
    while ((pos = s.find(em_dash, prev)) != string::npos) {
        if (pos > prev) {
            res.push_back(s.substr(prev, pos - prev));
        }
        prev = pos + em_dash.size();
    }
    if (prev < s.size()) {
        res.push_back(s.substr(prev));
    }
    return res;
}

int main() {
    ifstream file("example.txt");
    string token;
    while (file >> token) {
        vector<string> words = split_by_emdash(token);
        for (const auto& word : words) {
            cout << word << endl;
        }
    }
    return 0;
}

效果验证

针对你给出的测试输入He was young—perhaps from twenty-eight to thirty—tall, slender,代码输出结果完全符合预期:

He
was
young
perhaps
from
twenty-eight
to
thirty
tall,
slender

如果你使用宽字符流处理Unicode文件,只需将代码中string替换为wstring,长破折号常量改为L"\x2014"即可。

内容的提问来源于stack exchange,提问作者jaryl

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 08:36:02