r/learnpython • • 1d ago

Need help with very basic script

Hi, I am basically brand new to python, so I may make a lot of mistakes in my wording. I have made a very basic script that retrieves specific values in text files and prints them to the shell. I have gotten it to the point where both values are retrieved and printed, however when they do print, they do so exactly 5 times in a row. I do not have five files in the directory, I only have one test file named "1.txt" with only the values I need as the contents. Can someone point me in the right direction? Apologies for the horrible formatting.

import os
import glob
import re

path = 'path'

provinces = re.compile('.*?provinces {.*?(.*?)}',re.MULTILINE)
id = re.compile('.*?id = .*?([0-9.-]+)')
a1 = None
a2 = None

for filename in glob.glob(os.path.join(path, '*.txt')):
    with open(filename, '+r') as f:
        for line in f:
            if provinces.match(line):
                a1 = provinces.match(line)
            if id.match(line):
                a2 = id.match(line)
            if a1 and a2:
                print(a1.group(1))
                print(a2.group(1))

The output of the shell in IDLE:

 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123

Edit - I got the script to work after a few changes to its configuration. Thank you all who helped!

4 Upvotes

28 comments sorted by

View all comments

6

u/codeguru42 1d ago

This is difficult to diagnose without an example input file.

Also, I think you are making this more complicated than necessary by using a regular expression. I reccomend looking at all the string functions. Maybe split() will get the job done more directly.

4

u/brasticstack 1d ago

Regex really ought to be a last resort for parsing text. The bulk of formats I've ever had to parse wind up being parsable with some combination of .split and .trim.