You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python词频统计脚本中单词前'b'标识的成因及解决方法

Python词频统计输出出现b'xxx'标识的原因及解决方法

原因分析

urllib.request.urlopen() 获取到的响应内容是字节流(bytes类型),代码中直接对字节流的 line 调用 split() 得到的每个元素依然是bytes类型,Python在打印bytes对象时会自动加上b前缀来区分字节与字符串类型,这就是输出中出现b'yesterday,'的原因。

解决方法

将字节流解码为字符串即可,常用的编码格式是utf-8,可以在读取每行数据时调用decode()方法转换。

修改后的代码

import urllib.request

same_words = {}
text = "https://raw.githubusercontent.com/KseniaGiansar/pythonProject2_text/master/yesterday.txt"
request = urllib.request.urlopen(text)
each_word = []

# 收集单词到列表
for line in request:
    # 将字节流解码为字符串
    line_str = line.decode('utf-8')
    line_words = line_str.split()
    for word in line_words:
        each_word.append(word)

# 统计词频
for word in each_word:
    lower_word = word.lower()
    if lower_word not in same_words:
        same_words[lower_word] = 1
    else:
        same_words[lower_word] += 1

# 输出结果
for word, count in same_words.items():
    print(f"word = {word}, count = {count}")

关键修改点

  • 在处理每行数据时,新增line_str = line.decode('utf-8')将字节转为字符串
  • 移除了原代码中未使用的冗余变量,简化代码结构
  • 遍历字典时使用items()方法同时获取键值对,写法更高效

内容的提问来源于stack exchange,提问作者A_K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 05:45:37