You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除.txt文件换行符并存储到std::vector<std::string>中

问题描述

我有一个包含用户名及其ID的name_and_id.txt文件,内容如下:

Heal: A24-1234
Hael: A25-4567
Hela: A26-8910
Hale: A27-1112

Test1: A28-1314
Test2: A29-1516
Test3: A30-1718

我希望先去除所有换行符,再把文本存入std::vector<std::string>中。但我的代码在分词后输出仍有换行符,代码如下:

#include <algorithm>
#include <fstream>
#include <iostream>
#include <iterator>
#include <string>
#include <sstream>
#include <utility>
#include <vector>

auto main() -> int {
  
    std::fstream user_file{};
    
    const char space_char{' '}; //NOTE - space delimiter 
    std::string entries_by_line{};
    std::vector<std::string> data_entries{};
    
    std::stringstream entries_in_file{};
    
    std::string entries_data_imported_file{};
    std::vector<std::string> tokenized_data_imported_file{};
  
    user_file.open("name_and_id.txt");
    if(user_file.is_open()) {
        while(std::getline(user_file, entries_by_line)) { 
            data_entries.emplace_back(entries_by_line); //REVIEW - store data (without new lines) from `.txt` file into vector named `data_entries`
        }
  
        std::copy(
            data_entries.begin(),
            data_entries.end(),
            std::ostream_iterator<std::string>(entries_in_file, " ") //REVIEW - copy contents (with spaces) of vector `data_entries` to a std::stringstream named `entries_in_file`
        );
    
        std::cout << "Before removing punctuations:" << std::endl;
        for(size_t i{}; i != data_entries.size(); ++i) {
            std::cout << data_entries[i] << std::endl;
        }
  
        std::stringstream entries{std::move(entries_in_file.str())}; //REVIEW - store content of entries_in_file.str() to std::stringstream named `entries`
    
        while(std::getline(entries, entries_data_imported_file, space_char)) {
            entries_data_imported_file.erase(
                std::remove_if(
                    entries_data_imported_file.begin(),
                    entries_data_imported_file.end(),
                    ispunct //REVIEW - remove punctuations
                ),
                entries_data_imported_file.end()
            );
            tokenized_data_imported_file.emplace_back(entries_data_imported_file); //REVIEW - store data (without punctuations) to vector `tokenized_data_imported_file`
        }
        
    
        std::cout << "\nAfter removing punctuations:" << std::endl;
        for(size_t i{}; i != tokenized_data_imported_file.size(); ++i) {
            std::cout << tokenized_data_imported_file[i] << std::endl;
        }
    }
    user_file.close();
  
}

当前输出结果:

Before removing punctuations:
Heal: A24-1234
Hael: A25-4567
Hela: A26-8910
Hale: A27-1112

Test1: A28-1314
Test2: A29-1516
Test3: A30-1718

After removing punctuations and new line:
Heal
A241234
Hael
A254567
Hela
A268910
Hale
A271112

Test1
A281314
Test2
A291516
Test3
A301718

请问该如何去除输出中的换行符?


问题分析与解决

你的代码里,data_entries会把文件中的空行也存进去。因为std::getline遇到空行时,会读取到一个空字符串并存入vector,后续std::copy把这些空字符串也写入了entries_in_file,最终在分词时,空字符串会被保留,输出时就会出现空行。

核心解决步骤

  1. 读取文件时过滤空行
    在把读取到的行存入data_entries前,先判断该行是否为空,只保留非空的行。修改读取循环即可:
while(std::getline(user_file, entries_by_line)) { 
    // 跳过空行,只保留有效内容
    if (!entries_by_line.empty()) {
        data_entries.emplace_back(entries_by_line);
    }
}
  1. 可选:简化流程优化
    你当前的流程是「存行→转stringstream→分词」,其实可以直接读取每一行后拆分处理,避免中间冗余操作,同时彻底规避空行问题:
while(std::getline(user_file, entries_by_line)) { 
    if (entries_by_line.empty()) continue;
    // 按冒号拆分用户名和ID部分
    size_t colon_pos = entries_by_line.find(':');
    if (colon_pos != std::string::npos) {
        std::string name = entries_by_line.substr(0, colon_pos);
        // 处理ID:去掉空格和标点
        std::string id_part = entries_by_line.substr(colon_pos + 1);
        id_part.erase(std::remove_if(id_part.begin(), id_part.end(), [](char c) {
            return ispunct(c) || isspace(c);
        }), id_part.end());
        // 直接存入目标vector
        tokenized_data_imported_file.push_back(name);
        tokenized_data_imported_file.push_back(id_part);
    }
}

修改后输出效果

调整后,输出中的空行会完全消失,变为:

Before removing punctuations:
Heal: A24-1234
Hael: A25-4567
Hela: A26-8910
Hale: A27-1112
Test1: A28-1314
Test2: A29-1516
Test3: A30-1718

After removing punctuations:
Heal
A241234
Hael
A254567
Hela
A268910
Hale
A271112
Test1
A281314
Test2
A291516
Test3
A301718

内容的提问来源于Stack Exchange,提问作者Gil Rovero

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:41:08