You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python读取文本文件并识别指定标记?代码逻辑问题求助

词法分析器问题修复

问题背景

编写Python程序读取文本文件并标记语法组件时,出现逻辑错误:程序逐字符处理内容,无法识别完整的token(如把"input"拆分为单个字符分别标记),导致输出不符合预期。

现有代码

#open a text file
program = open("input.txt", "r");
strHold =""
x=""

#function to check the strings
def findWord(strIn):
    if(strIn == "input"):
        print("<input>, "+strIn)
        # strIn = ""
        return
        
    elif(strIn == "("):
        print("<lparen>, "+strIn)
        # strIn = ""
        return
    
    elif(strIn == ")"):
        print("<rparen>, "+strIn)
        # strIn = ""
        return
        
    elif(strIn == "="):
        print("<assign_op>, "+strIn)
        # strIn = ""
        return
        
    elif(strIn == "+" or strIn == "-"):
        print("<add_op>, "+strIn)
        # strIn = ""
        return
        
    elif(strIn == "+" or strIn == "-"):
        print("<add_op>, "+strIn)
        # strIn = ""
        return
        
    elif(strIn=="/" or strIn=="*" or strIn == "//" or strIn == "%"):
        print("<mult_op>, "+strIn)
        # strIn = ""
        return
    
    elif(strIn==" " or strIn=="\n"):#check
        x= strIn.isspace()
        # if(x==True):
            # strIn=""
        return
    
    elif(strIn.isnumeric() == True):
        print("<number>, "+strIn)
        # strIn=""
        return
    
    elif(strIn=="output"):
        print("<output>, "+strIn)
        #strIn = ""
        return
    
        
    else:#default is an id??
        print("<id>, "+strIn)
        # strIn=""
        return


#loop through .txt file
for line in program:
    for c in line:
        strHold = strHold+c
        findWord(strHold);
        strHold="" 

输入文件(input.txt)

input(a)
input(b)
input(c)
total = a + b + c  /* get a sum of three inputs */
average = total / 3 /* compute an average */
output(total)
output(average)

当前错误输出

<id>, i
<id>, n
<id>, p
<id>, u
<id>, t
<lparen>, (
<id>, a
<rparen>, )
<id>, i
...

期望输出

<input>, input
<lparen>, (
<id>, a
<rparen>, )

修复方案

核心问题是原代码每读取一个字符就立即调用识别函数并清空缓冲区,完全没有机会拼接成完整的token。需要调整逻辑:先将字符拼接成完整的token,直到遇到分隔符、运算符或匹配到关键字时,再进行识别并重置缓冲区。

具体修改要点

  1. 重构字符读取逻辑:不再逐字符处理后立即清空缓冲区,而是持续拼接字符,直到确定token结束。
  2. 处理多字符运算符:比如//这类双字符运算符,需要检查后续字符是否构成完整运算符。
  3. 跳过注释内容:识别并跳过/* ... */格式的注释,避免错误标记。
  4. 移除无效代码:原函数内的strIn = ""是局部变量赋值,无法改变外部缓冲区,直接删除。

修改后的代码

def findWord(strIn):
    if strIn == "input":
        print("<input>,", strIn)
    elif strIn == "(":
        print("<lparen>,", strIn)
    elif strIn == ")":
        print("<rparen>,", strIn)
    elif strIn == "=":
        print("<assign_op>,", strIn)
    elif strIn in ("+", "-"):
        print("<add_op>,", strIn)
    elif strIn in ("/", "*", "//", "%"):
        print("<mult_op>,", strIn)
    elif strIn.isnumeric():
        print("<number>,", strIn)
    elif strIn == "output":
        print("<output>,", strIn)
    elif strIn.strip() == "":
        pass
    else:
        print("<id>,", strIn)

def tokenize_line(line):
    strHold = ""
    i = 0
    n = len(line)
    in_comment = False
    while i < n:
        c = line[i]
        # 处理注释块
        if in_comment:
            if c == "*" and i+1 < n and line[i+1] == "/":
                in_comment = False
                i += 2
            else:
                i += 1
            continue
        
        if c == "/" and i+1 < n and line[i+1] == "*":
            in_comment = True
            i += 2
            continue
        
        # 处理空白字符:先识别缓冲区内容,再跳过空白
        if c.isspace():
            if strHold:
                findWord(strHold)
                strHold = ""
            i += 1
            continue
        
        # 处理单字符分隔符/运算符
        if c in ("(", ")", "=", "+", "-", "%"):
            if strHold:
                findWord(strHold)
                strHold = ""
            findWord(c)
            i += 1
            continue
        
        # 处理多字符运算符
        if c == "/" and i+1 < n and line[i+1] == "/":
            if strHold:
                findWord(strHold)
                strHold = ""
            findWord("//")
            i += 2
            continue
        if c == "*":
            if strHold:
                findWord(strHold)
                strHold = ""
            findWord("*")
            i += 1
            continue
        
        # 拼接字母/数字,匹配关键字时立即识别
        if c.isalnum():
            strHold += c
            if strHold in ("input", "output"):
                findWord(strHold)
                strHold = ""
            i += 1
        else:
            # 处理未知字符
            if strHold:
                findWord(strHold)
                strHold = ""
            findWord(c)
            i += 1
    # 处理行尾剩余的token
    if strHold:
        findWord(strHold)

# 安全读取文件并处理每一行
with open("input.txt", "r") as program:
    for line in program:
        tokenize_line(line)

代码说明

  • 使用with语句管理文件,自动关闭文件,避免资源泄漏。
  • 新增tokenize_line函数专门处理整行的token拆分,解决逐字符处理的核心问题。
  • 加入注释识别逻辑,跳过注释内容。
  • 优先处理多字符运算符和关键字,确保完整识别。
  • 仅在遇到分隔符、运算符或匹配到关键字时,才识别缓冲区内容并重置。

内容的提问来源于stack exchange,提问作者brocoli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 02:05:34