You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:如何拆分列表中的字符串并清理特殊字符?

解决字符串列表拆分与清理问题

Got it, let's work through this problem together. You need to take a list of strings containing newlines (\n), form feeds (\x0c), and spaces, then turn them into a flat list of clean individual words. Here's a straightforward, reliable way to do this in Python:

Step-by-Step Solution

The core idea is to treat all whitespace-like characters (spaces, newlines, form feeds) as separators, while making sure we don't end up with empty strings in the final result.

Option 1: Explicit Loop (Easy to Follow)

import re

original_list = ["this is an\nexample" , "sentence\n\x0c"]
clean_words = []

for string in original_list:
    # Strip leading/trailing whitespace (including \n and \x0c) first
    stripped_str = string.strip()
    # Split on one or more whitespace characters (handles all types of whitespace)
    split_words = re.split(r'\s+', stripped_str)
    # Add the valid words to our result list
    clean_words.extend(split_words)

print(clean_words)  # Output: ["this", "is", "an", "example", "sentence"]

Option 2: Concise List Comprehension

If you prefer a more compact approach, a nested list comprehension achieves the same goal:

import re

original_list = ["this is an\nexample" , "sentence\n\x0c"]
clean_words = [word for s in original_list for word in re.split(r'\s+', s.strip()) if word]

print(clean_words)

Why This Works

  • string.strip(): Removes any leading or trailing whitespace characters (including \n and \x0c) from each string, eliminating empty strings that might come from splitting strings that start/end with these characters.
  • re.split(r'\s+', ...): Uses a regular expression to split the string on one or more whitespace characters. The \s pattern matches all whitespace types (spaces, tabs, newlines, form feeds), and the + ensures consecutive whitespace is treated as a single separator.
  • The if word check (in the list comprehension) is an extra safeguard to filter out any stray empty strings, though strip() combined with \s+ should already prevent this.

内容的提问来源于stack exchange,提问作者EzekielGMG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:26:42