Python3在文本特定索引位置插入标签导致单词截断的问题求助
I see the issue here—when you insert tags into the character list, each insertion shifts the indices of all characters that come after it. Since you're processing labels from first to last, the original fin positions you're using are no longer accurate once you've added tags earlier in the text. That's why your tags end up cutting off mid-word.
The Solution: Process Labels in Reverse Order
The fix is straightforward: handle your labels starting from the last one and moving to the first. This way, inserting tags at the end of the text doesn't affect the indices of labels that appear earlier (since those are in the unmodified portion of the character list).
Modified Code
texte = { "text": "The applicant's cells were overcrowded. The detainees had to take turns to sleep because there was usually one sleeping place for two to three of them. There was almost no light in the cells because of the metal shutters on the windows, as well as no fresh air. The lack of air was aggravated by the detainees' smoking and the applicant, a non-smoker, became a passive smoker. There was one hour of daily exercise. The applicant's eyesight deteriorated and he developed respiratory problems. In summer the average air temperature was around thirty degrees which, combined with the high humidity level, caused skin conditions to develop. The sanitary conditions were below any reasonable standard. In particular, the cells were supplied with water for only one or two hours a day and on some days there was no water supply at all. The lack of water caused intestinal problems and in 1999 the administration had to announce quarantine in that connection.", "label": [ [328,347,"Article 3 - Violated"], [2269,2323,"Article 3 - Violated"], [2791,2843,"Article 3 - Violated"], [2947,2988,"Article 3 - Violated"], [3099,3110,"Article 3 - Violated"], [3603,3615,"Article 3 - Violated"], [3702,3756,"Article 3 - Violated"], [4793,4923,"Article 3 - Violated"], [5185,5196,"Article 3 - Violated"], [8111,8198,"Article 3 - Respected"], [8510,8521,"Article 3 - Respected"], [8575,8601,"Article 3 - Respected"], [8965,9009,"Article 3 - Respected"] ] } # Convert text to a mutable character list text_chars = list(texte["text"].strip()) labels = texte["label"] # Process labels in reverse to avoid index shifting errors for label in reversed(labels): start_idx = label[0] end_idx = label[1] tag_name = label[2] # Insert closing tag right after the end of the target segment text_chars.insert(end_idx + 1, f"</{tag_name}>") # Insert opening tag at the start of the target segment text_chars.insert(start_idx, f"<{tag_name}>") # Join the character list back into a single string formatted_text = ''.join(text_chars) print(formatted_text)
Key Explanations
- Reverse Processing: By iterating over labels from last to first, we ensure that inserting tags for later text segments doesn't shift the indices of earlier segments. This keeps your original
start_idxandend_idxvalues accurate relative to the original text. - Tag Placement: We insert the closing tag at
end_idx + 1to place it immediately after the last character of the target segment, and the opening tag atstart_idxto place it right before the first character. - Clean Join: Using
''.join(text_chars)efficiently converts the modified character list back into a single string without any extra spaces or artifacts.
Quick Check
Double-check that your end_idx values correctly point to the last character of the segment you want to tag. If the index is off by one (e.g., pointing to the character before the last one), that could still cause truncation—but the reverse processing fix addresses the core index-shifting issue.
内容的提问来源于stack exchange,提问作者Balkhrod

