循环迭代DataFrame列时创建动态命名的独立列表
现有一个包含多列的DataFrame,需遍历指定列,通过名为tokenizer的函数将每一行转换为列表。相关代码如下:
columns = ['stemmed', 'lemmatized', 'lem_stop','stem_stop', 'lem_stop_nltk', 'stem_stop_nltk'] tokenizer = nltk.data.load('tokenizers/punkt/english.pickle') sentences = [] for i in columns: for tweet in df[i]: sentences += tweet_to_sentences(tweet, tokenizer)目前所有结果均存入同一个sentences列表,现需为每个列创建一个独立的列表,而非合并存储,请教实现方法。
Hey, this is a common scenario—instead of dumping everything into one big list, you can use a dictionary to map each column name to its own dedicated list of sentences. This makes it super easy to access each column's results individually. Here are two straightforward ways to implement this:
Method 1: Initialize an empty dictionary and populate column by column
This approach is easy to read and debug, which is great if you might need to add extra logic later:
columns = ['stemmed', 'lemmatized', 'lem_stop','stem_stop', 'lem_stop_nltk', 'stem_stop_nltk'] tokenizer = nltk.data.load('tokenizers/punkt/english.pickle') # Create an empty dict where keys are column names, values are their sentence lists column_sentences = {} for col in columns: # Start with an empty list for the current column column_sentences[col] = [] for tweet in df[col]: # Add all processed sentences from this tweet to the column's list column_sentences[col].extend(tweet_to_sentences(tweet, tokenizer))
After running this, you can access each column's sentences directly using the column name as a key—like column_sentences['stemmed'] or column_sentences['lem_stop_nltk'].
Method 2: Use a dictionary comprehension (more concise)
If you prefer shorter, more Pythonic code, a dictionary comprehension does the exact same job in a compact way:
columns = ['stemmed', 'lemmatized', 'lem_stop','stem_stop', 'lem_stop_nltk', 'stem_stop_nltk'] tokenizer = nltk.data.load('tokenizers/punkt/english.pickle') column_sentences = { col: [sentence for tweet in df[col] for sentence in tweet_to_sentences(tweet, tokenizer)] for col in columns }
This works exactly like the first method, just wraps the loop logic into a single expression.
Quick tip
If you ever need to extract a single column's list into its own variable, just do:
stemmed_sentences = column_sentences['stemmed']
内容的提问来源于stack exchange,提问作者A.Papa

