如何去除.txt文件换行符并存储到std::vector<std::string>中
问题描述
我有一个包含用户名及其ID的name_and_id.txt文件,内容如下:
Heal: A24-1234 Hael: A25-4567 Hela: A26-8910 Hale: A27-1112 Test1: A28-1314 Test2: A29-1516 Test3: A30-1718
我希望先去除所有换行符,再把文本存入std::vector<std::string>中。但我的代码在分词后输出仍有换行符,代码如下:
#include <algorithm> #include <fstream> #include <iostream> #include <iterator> #include <string> #include <sstream> #include <utility> #include <vector> auto main() -> int { std::fstream user_file{}; const char space_char{' '}; //NOTE - space delimiter std::string entries_by_line{}; std::vector<std::string> data_entries{}; std::stringstream entries_in_file{}; std::string entries_data_imported_file{}; std::vector<std::string> tokenized_data_imported_file{}; user_file.open("name_and_id.txt"); if(user_file.is_open()) { while(std::getline(user_file, entries_by_line)) { data_entries.emplace_back(entries_by_line); //REVIEW - store data (without new lines) from `.txt` file into vector named `data_entries` } std::copy( data_entries.begin(), data_entries.end(), std::ostream_iterator<std::string>(entries_in_file, " ") //REVIEW - copy contents (with spaces) of vector `data_entries` to a std::stringstream named `entries_in_file` ); std::cout << "Before removing punctuations:" << std::endl; for(size_t i{}; i != data_entries.size(); ++i) { std::cout << data_entries[i] << std::endl; } std::stringstream entries{std::move(entries_in_file.str())}; //REVIEW - store content of entries_in_file.str() to std::stringstream named `entries` while(std::getline(entries, entries_data_imported_file, space_char)) { entries_data_imported_file.erase( std::remove_if( entries_data_imported_file.begin(), entries_data_imported_file.end(), ispunct //REVIEW - remove punctuations ), entries_data_imported_file.end() ); tokenized_data_imported_file.emplace_back(entries_data_imported_file); //REVIEW - store data (without punctuations) to vector `tokenized_data_imported_file` } std::cout << "\nAfter removing punctuations:" << std::endl; for(size_t i{}; i != tokenized_data_imported_file.size(); ++i) { std::cout << tokenized_data_imported_file[i] << std::endl; } } user_file.close(); }
当前输出结果:
Before removing punctuations: Heal: A24-1234 Hael: A25-4567 Hela: A26-8910 Hale: A27-1112 Test1: A28-1314 Test2: A29-1516 Test3: A30-1718 After removing punctuations and new line: Heal A241234 Hael A254567 Hela A268910 Hale A271112 Test1 A281314 Test2 A291516 Test3 A301718
请问该如何去除输出中的换行符?
问题分析与解决
你的代码里,data_entries会把文件中的空行也存进去。因为std::getline遇到空行时,会读取到一个空字符串并存入vector,后续std::copy把这些空字符串也写入了entries_in_file,最终在分词时,空字符串会被保留,输出时就会出现空行。
核心解决步骤
- 读取文件时过滤空行
在把读取到的行存入data_entries前,先判断该行是否为空,只保留非空的行。修改读取循环即可:
while(std::getline(user_file, entries_by_line)) { // 跳过空行,只保留有效内容 if (!entries_by_line.empty()) { data_entries.emplace_back(entries_by_line); } }
- 可选:简化流程优化
你当前的流程是「存行→转stringstream→分词」,其实可以直接读取每一行后拆分处理,避免中间冗余操作,同时彻底规避空行问题:
while(std::getline(user_file, entries_by_line)) { if (entries_by_line.empty()) continue; // 按冒号拆分用户名和ID部分 size_t colon_pos = entries_by_line.find(':'); if (colon_pos != std::string::npos) { std::string name = entries_by_line.substr(0, colon_pos); // 处理ID:去掉空格和标点 std::string id_part = entries_by_line.substr(colon_pos + 1); id_part.erase(std::remove_if(id_part.begin(), id_part.end(), [](char c) { return ispunct(c) || isspace(c); }), id_part.end()); // 直接存入目标vector tokenized_data_imported_file.push_back(name); tokenized_data_imported_file.push_back(id_part); } }
修改后输出效果
调整后,输出中的空行会完全消失,变为:
Before removing punctuations: Heal: A24-1234 Hael: A25-4567 Hela: A26-8910 Hale: A27-1112 Test1: A28-1314 Test2: A29-1516 Test3: A30-1718 After removing punctuations: Heal A241234 Hael A254567 Hela A268910 Hale A271112 Test1 A281314 Test2 A291516 Test3 A301718
内容的提问来源于Stack Exchange,提问作者Gil Rovero
相关产品推荐
相关产品推荐

