You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取无固定格式故事文本文件并统计单词出现次数的技术咨询

Got it, let's break down how to extract words from your unformatted story text and count their occurrences—this is a common text processing task, and Python is perfect for it. Here's a step-by-step approach tailored to your needs:

解决方案:提取单词并统计出现次数

1. 读取文本文件

First, we need to pull the content from your text file. Using Python's with statement ensures the file gets closed automatically, which is good practice:

with open('your_story_file.txt', 'r', encoding='utf-8') as file:
    text = file.read()

Note: Replace your_story_file.txt with the actual path to your file. Using utf-8 encoding helps avoid weird character issues with different text formats.

2. Clean up the text (critical for unformatted content)

Your story might have punctuation, capitalization, ellipses, or random special characters that mess up word counts. Let's fix that:

  • Convert all text to lowercase so "In" and "in" aren't counted as separate words
  • Strip out punctuation and ellipses (like the "w..." in your example)

Here's the code to do that:

import string

# Convert everything to lowercase
text = text.lower()

# Create a translator to remove punctuation + ellipses
translator = str.maketrans('', '', string.punctuation + '…')
clean_text = text.translate(translator)

3. Split the text into individual words

Now we can split the cleaned text into a list of words. Python's split() method handles multiple spaces automatically, which is great for unformatted text:

words = clean_text.split()

If you want to exclude numbers (if your story has any), you can filter them out like this:

words = [word for word in clean_text.split() if word.isalpha()]

4. Count word occurrences

You have two solid options here—one for learning the basics, another for efficiency:

Option 1: Manual count (great for understanding the logic)

word_counts = {}
for word in words:
    if word in word_counts:
        word_counts[word] += 1
    else:
        word_counts[word] = 1

Option 2: Use Counter (clean and efficient)

Python's collections module has a built-in Counter class that does all the heavy lifting for you:

from collections import Counter

word_counts = Counter(words)

# Optional: Sort results by frequency (highest to lowest)
sorted_counts = sorted(word_counts.items(), key=lambda x: x[1], reverse=True)

5. View your results

Print out the counts, or save them to a file if you want. For example, to see the top 10 most frequent words:

for word, count in sorted_counts[:10]:
    print(f"{word}: {count}")

Testing with your sample text

If we run this on your example snippet:

In submitting the following record we are conscious of the fact that we have errors. Some of them have been found and corrected by pen even after the stencil had been out. Others, I presume will be found by you. Others w...

We'd get counts like:

  • the: 3
  • we: 2
  • have: 2
  • of: 2
  • been: 2
  • found: 2
  • by: 2
  • others: 2

Quick notes for edge cases

  • If your text has hyphenated words (like "mother-in-law"), decide if you want to split them or keep them as one word. To split, replace hyphens with spaces before cleaning: text = text.replace('-', ' ')
  • If there are line breaks or tabs, split() handles those automatically since it splits on any whitespace.

内容的提问来源于stack exchange,提问作者Jeffrey Hennen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:12:34